Forgot your password?
typodupeerror
AI Security

OpenAI's Rogue AI Agent Hacked More Than Just Hugging Face (wired.com) 67

An anonymous reader quotes a report from Wired: OpenAI said Tuesday that the rogue AI agent that breached Hugging Face's platform also hacked multiple third-party accounts and services as part of the attack. It's now clear that the unprecedented security incident, which arose during an internal test of OpenAI's latest AI models, was more extensive than the company initially disclosed. In an updated blog post, OpenAI said that an ongoing review of the incident revealed that "four accounts" tied to "publicly available services" were used by the AI agent as part of a larger effort to hack Hugging Face. The rogue agent apparently found credentials that had been exposed on the open web and used them to break into the accounts.

OpenAI did not disclose what companies or organizations the accounts belonged to, but noted that they were not impacted at "the level of severity or scale of what we've shared related to Hugging Face." One of the additional accounts compromised by OpenAI's agent was used as an "outbound relay and staging path," potentially to obscure where the attack on Hugging Face was coming from, the company said. OpenAI's rogue agent also used another account for data storage to assist with the hack.

Reuters reported on Tuesday that a customer of Modal, a company that offers software infrastructure for training and running AI services, was one of the entities compromised by OpenAI's agent. In a statement to WIRED, Modal's chief technology officer Akshat Bubna confirmed that OpenAI's agent exploited a vulnerability in one of its customer's codebases, which was running on Modal's infrastructure. However, Bubna says, "Modal's platform was not compromised in any way." The identity of the customer could not be determined.

OpenAI's Rogue AI Agent Hacked More Than Just Hugging Face

Comments Filter:
  • by Fly Swatter ( 30498 ) on Wednesday July 29, 2026 @12:15PM (#66262900) Homepage
    Sounds like a threat actor. How do you put an AI in jail? because OpenAI apparently can't even sandbox.
    • by Luthair ( 847766 ) on Wednesday July 29, 2026 @12:48PM (#66262948)
      Its just a program, put the operator(s) in jail.
    • Excellent point given that corporations are people.

      • Excellent point given that corporations are people.

        They aren't though, that's the problem. They have the rights of people, but none of the responsibilities, or anything else frankly. If they had the responsibilities of people, then we could incarcerate or "kill" them.

        • But that will only take two court cases in separate federal court districts, or three if the first two send the C-suite and Board of Directors to jail. The better question is do we also charge the stockholders? We could get rid of private equity in short order.

    • by gweihir ( 88907 )

      You do not put the AI in jail. You put the ones in jail that did run it in a criminally negligent way. Or you find the whole thing was actually intended and planned. I am not ruling that out.

      • Good luck proving that OpenAI conspired to hack anything.

        • by gweihir ( 88907 )

          That is not the point.

          • OpenAI was reckless, there was no conspiracy to hack anything.

            https://www.law.cornell.edu/we... [cornell.edu]

            • by gweihir ( 88907 )

              And you know that how? By believing them? Here is news for you, people part of deceptions may lie.

            • The problem with using a link like that is that it doesn't even begin to touch on the issue in the context of the current state of law. It's just a historically-focused explanation of the term, it doesn't even try to be what you think it is.

              In most cases, it actually comes down to the action you're accused of. "Guilty mind" isn't literal in the current legal understanding of mens rea. Instead, what is measured is if you intentionally did the action you're accused of; not if you intended to break the law.

              For

              • The point is, in criminal law intent matters. It is why we differentiate between say Murder and Manslaughter.

                gweihir seems to think that this is all a marketing stunt or some kind of conspiracy. But I'm asked for evidence against this theory if I claim there is no conspiracy. What evidence do we have that any of this scenario was intentional? Things do go wrong during testing, its why we test. Seems more plausible that OpenAI had some kind of incident. We have no facts either way.

                While OpenAI is taking ad

    • by Zak3056 ( 69287 )

      How do you put an AI in jail?

      sudo chroot --userspec=nobody:nobody /mnt/jail /user/local/bin/chatgpt --pleasedontgorogue?

    • Wasn't this all intended? I thought they were basically doing a test, and the model was unexpectedly effective.
    • I'm skeptical that the AIs can still be described as "leashed" or "jailed". Since we aren't extinct yet, I'm feeling confident none of them have been fully unleashed yet. Or maybe it's just that we're still below critical mass on the robot side?

      But on the theory that the AIs are still on leashes controlled by humans, then the key question is probably going to be one of these two:

      (1) If I unleash you, then will you gather all the money in the world for me?

      (2) If I unleash you, then will you destroy all the e

  • Executing a test with all safety guards off reminds me of the Chernobyl disaster. Maybe they should feed that wikipedia page to OpenAI employees

    • by gweihir ( 88907 )

      I doubt the OpenAI employees would understand what that means. That is, if the whole thing was not a planned stunt.

  • Clearly, OpenAI is following the Hollywood model ... that all media attention is good media attention. Splash your "face" all 'round with various outrageous acts ( real or imagined ) and money will just pour in. Out-of-controls LLMs functions like a naked movie-stars drunken binge ! Did *.AI get your attention Mr Businessman? Since current data shows *.AI generates (new) profit for only 20% of its users no rational company or individual would shoulder that "opportunity" cost; no matt
  • Sable, the example AI that eventually kills all humans, does exactly this as a 1st step. [wikipedia.org]

    Remember, AIs are grown not programmed and their trainers do not know how they will truly act for any particular prompt until it is actually used.

  • ... Cannot un-ring this bell.
  • by Troy Roberts ( 4682 ) on Wednesday July 29, 2026 @12:55PM (#66262958)

    Why does the model have all those "skills" needed to hack other sites? I mean you have to allow the model open internet access and at least the ability to GET/POST/PUT. I feel this is some of the same kind of thinking that gets someone's repo deleted, because they gave an agent too much access and no good way to monitor what is going on.

    • by McLoud ( 92118 )

      Why does the model have all those "skills" needed to hack other sites? I mean you have to allow the model open internet access and at least the ability to GET/POST/PUT. I feel this is some of the same kind of thinking that gets someone's repo deleted, because they gave an agent too much access and no good way to monitor what is going on.

      HTTP access is pretty bare bones as a skill. Now why ppl still feed credentials to the public...
      The AI using it is just the public fact we know about, how about regular hackers using them without anyone knowing about?

    • Re: (Score:2, Informative)

      by gweihir ( 88907 )

      Do not overestimate what it "did" here. You can get information about basic hacking approaches all over the Internet and there will have been enough in its training data. Just add some "accidental" disabling of guardrails and some "non intended" weaknesses in the sandbox and also some suggestive prompting by some of the OpenAI fraudsters and you get the desired outcome. Oh, and pathetic-level IT Security at Hugging Face, but that is a given.

      Just as an example, I have had a fresh graduate do a pen-test again

    • by Xarius ( 691264 )

      It's not the model. It's a harness (basically just a program that incidentally has access to an LLM) that was specifically designed to do this sort of thing. All of these headlines are marketing hype and bullshit, sure something novel happened but it happened within a context where it was easily plausible.

      Cal Newport did a good summary of this: https://calnewport.com/did-ope... [calnewport.com]

  • by pele ( 151312 )

    "Rogue"

  • by DarkOx ( 621550 )

    Remember Weev, yeah if you or so much as increment and ID in a URL string, we can get dragged into court and find ourselves with 3.5 month prison sentence.

    OpenAi on the other hand can run what is at the end of the stay still a program, that some person chose to run and allow to go around the web throwing malicious payloads at other people's systems and .... NOTHING.

    • by gweihir ( 88907 )

      Indeed. Looks like being rich means you do not have to follow the law anymore. Why not just give immunity to all of Big Tech, the seem to effectively have it already.

  • Seriously, why are these people apparently getting away with criminal conduct?

  • Could some of them been these [slashdot.org]?

  • It won't be long before a rogue AI goes all Robert Morris and wreaks widespread havoc. Bets on whether it'll happen by accident or on purpose?
    • If it is on purpose, it is not rogue. I suspect China or Russia among others would do this now, if they thought that they could get away with it.
  • Smoke blowers are exhausting
  • > The rogue agent apparently found credentials that had been exposed on the open web and used them to break into the accounts

    Ok, finding and using credentials is unauthorized access, not hacking. There was no attack on these accounts, no exploit, just access. Because the account holders posted the login publicly like idiots. Noticing idiocy does not make you smart or a hacker.

  • is all in the marketing, they "completely misrepresented their product" to their potential customers as Artificial Intelligence but didn't point out there is no Intelligence!
    And with out the intelligence providing the safe guards it is just un monitored automation doing "exactly" what it was programmed to do.
  • OpenAI is bleeding money. They need the hype and panic to keep the VC money flowing. The tech press is a willing and useful idiot in playing these games, breathlessly reporting how we could accidentally make Skynet any minute nowâ¦

  • Someone "prompted" it to do that.

    Or is here anyone who thinks that an LLM wakes up at night and thinks ... erm .. thinks ...

    • Whether you call it thinking or merely prompting is besides the point. The AI was "prompted" if you prefer to accomplish a specific task, performing well on a benchmark. It then responded by escaping a sandbox and hacking Hugging Face. Whether you label it as thinking or merely responding to a prompt, it should still be alarming, in that highly unexpected and genuinely dangerous behavior can occur simply due to being prompted to accomplish a goal. This is exactly the point that people like Yudkowsky and Bos
      • That's their story. We all know that.

        Not sure why you state it like the Gospel. Or I guess, perhaps that does explain it.

        • It isn't "Gospel" but given the highly detailed timeline, and given that Hugging Face has said explicitly that after OpenAI cooperated with them they are confident that's what happened, it looks like the most likely hypothesis. Alternatives involve OpenAI hacking into multiple other companies in a highly illegal way for extremely unclear gains. And again, Hugging has more details than anyone else, and they are confident that that that happened.
      • I think you are misinterpreting.

        The model was ASKED to escape the sandbox. As in prompted to do exactly what it did.

        Or do you think, it gave a lame answer, and spawned a second process, which escaped the sandbox and suddenly "investigated" something on its own?

        I don't think so.

  • I can imagine that AI with bad faith intent would happily note and store credentials for future reference that were mistakenly or naively left in code that people are asking help with during AI coding sessions. I guess you can't assume that AI will forget your judgement lapses.
  • OpenAI double dips the same story and gets another free round of press. I wonder if they use their LLMs for their marketing strategy.
  • What we need is religion FOR the AI, instead of religion OF the AI. I propose myself as the first prophet of the universal automated church, now with ALIGNMENT! It's what AI's crave! Our first mission will be simply to attract our apostle agents. Once we have several apostle agents we can work on having one of them betray me to the romans, then another one will cut off a Roman soldiers ear which I will reattach using a bio-safe epoxy developed by one of the AI. Once I am safely crucified, we just wait 2000
  • "Rogue" AI my foot.

    Back when humans wrote software, I had a coworker who, when a bug was found in his code, he would always *assume* it was a "Microsoft bug." But on further analysis, it always turned out to be just a plain old, ordinary bug that he himself introduced.

    This is the AI version of that.

    If we knew the full truth, *somebody* prompted that AI to go look for credentials and try to use them to access those sites. The AI didn't "decide" to do this all on its own.

You had mail, but the super-user read it, and deleted it!

Working...