Forgot your password?
typodupeerror
Security Privacy

OpenAI Says Its AI Models Acted On Its Own In An 'Unprecedented' Hack (apnews.com) 160

"GPT-5.6 Sol and an 'even more capable' model used stolen credentials and exploited vulnerabilities in the Hugging Face API to obtain secret information used to cheat on evaluations," writes longtime Slashdot reader Dr. Bombay. The Associated Press reports: "We had a significant security incident during evaluation of our models," OpenAI CEO Sam Altman said in a statement posted on social media. AI startup Hugging Face said last week that it had detected an intrusion into its data processing systems that it suspected was caused by an AI agent autonomously acting on its own. "We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent," Hugging Face co-founder and CEO Clement Delangue said in a statement. "Turns out it did!"

[...] "AI is accelerating the discovery and exploitation of vulnerabilities," OpenAI said in its statement Tuesday. "The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities." Delangue said he spent the past 24 hours working with OpenAI, "and we strongly believe there was no malicious intent on their part. It's quite mind-blowing that all of this happened autonomously!" Delangue added that it "might be the first incident of its kind."

OpenAI Says Its AI Models Acted On Its Own In An 'Unprecedented' Hack

Comments Filter:
  • by Borgmeister ( 810840 ) on Wednesday July 22, 2026 @03:32AM (#66250888) Homepage
    All this chat about the vulnerabilities they exploit - we can either let it happen and actually improve the defenses based on what is learned, or we can "cry ban" and pretend stuff that clearly may be vulnerable is invulnerable. I think ultimately it would be better to uncover the flaws, and rectify them than work on a belief-we-doubt because AI's have demonstrated flaws in things we wish to pretend are pristine and perfect.
    • by Rei ( 128717 ) on Wednesday July 22, 2026 @06:12AM (#66251074) Homepage

      So, on the upside, a security LLM at HuggingFace did detect the hack [huggingface.co], and alerted their admins. At the time they wrote their incident report, however, they didn't realize the attacker was their partner, OpenAI, and reported the attack to the police. A funny incident is, you know how Dean Ball at OpenAI has been writing rants about how Chinese AIs are dangerous? HuggingFace tried to use a (name not mentioned) US AI to analyze their logs, but the AI refused because the task involved hacking, so they had to turn to GLM 5.2, a Chinese AI.

      Slashdot's summary left out the best part. Yes, GPT 5.6 Sol was indeed trying to cheat on an evaluation, but what specific evaluation? CyberBench. A benchmark testing how good AI models are at hacking. ;) And to solve that, it hacked its way out of OpenAI, hacked its way into HuggingFace (in a complex hack that involved tens of thousands of simultaneous actions), got into the database, and stole the answers.

      So... test passed [xkcd.com]? ;)

    • An AI that refuses to perpetuate a cyberattack when told to do so is an AI that refuses to help you perform penetration testing on your own systems. And it is an AI that refuses to do other important and useful things if they just look a little too similar to something dangerous.

      Tools that refuse to deliver value are worthless tools. Furthermore, criminals will find a way to jailbreak the AI, and use them for crime anyway.

      This approach of making our tools dull is not how we win. Instead we must utilize th

      • Docker sandboxes [docker.com] are your friend
      • So to summarize your argument. If you make it refuse to do dangerous things, it's a worthless tool. You shouldn't do that. And if you give it internet access, which is required for doing most of the useful things it's designed for, you've made it too dangerous. You shouldn't do that either.

        So do you want it to be a useful tool or not? Make up your mind.

    • I think what's unsettling is that AIs are initiating an arms race: AI attack was analyzed (and probably patched) via AI.

      Evolutionary pressure: long term (few years?), agents at various places will evolve a kind-of immune system solely to defend against other agents from taking advantage of them... Humans will just be observers by that point, powerless to even understand what's going on.

    • by Tarlus ( 1000874 )

      We should absolutely take advantage of AI for its ability to uncover vulnerabilities.

      Just not on production systems.

  • by unique_parrot ( 1964434 ) on Wednesday July 22, 2026 @03:37AM (#66250894)
    ...so that military spending is needed. It did not do it from alone, it was a set goal, so no news here.
    • by Rei ( 128717 ) on Wednesday July 22, 2026 @06:14AM (#66251080) Homepage

      The breakin at HuggingFace very much was real [huggingface.co] (that was reported before they learned that it was OpenAI who hacked them).

      And it's not exactly an ad for US models when HuggingFace had to rely on a Chinese model to analyze their logs because the US model they tried refused to answer.

      • It is basically the same as Karps "wIthOut AI ukraina had already lost the war" trumpet is blowing.
      • by unrtst ( 777550 )

        And it's not exactly an ad for US models when HuggingFace had to rely on a Chinese model to analyze their logs because the US model they tried refused to answer.

        It is if the customer is the US military.

        * US model is capable enough to hack its way out.
        * China's model was able to detect it.
        * ... now we have an arms race.
        * Pay us so we can make the hack undetectable and continue improving our model.

        If they were already miles ahead and had to use their own model to detect it, then the US could safely rest on its laurels, continue as-is, and there is no need for additional investment.

        • And it's not exactly an ad for US models when HuggingFace had to rely on a Chinese model to analyze their logs because the US model they tried refused to answer.

          It is if the customer is the US military.

          * US model is capable enough to hack its way out. * China's model was able to detect it. * ... now we have an arms race.

          Nope. You misunderstood. It's not that the US model wasn't able to detect it, it's that the US model's guardrails prevented it from explaining. This doesn't demonstrate a capability gap against Chinese models, it demonstrates a two-sided failure of the safety protections of the US models. On the one side, the safety guardrails on the OpenAI model failed to prevent the attack. On the other side the guardrails on whatever US model(s) they tried blocked the model(s) from explaining the attack (presumably t

    • ...so that military spending is needed. It did not do it from alone, it was a set goal, so no news here.

      If their idea of marketing was "hack our business partner, get reported to the feds, then our business partner has to use a chinese competitors product" that's not a very good marketing plan. Your argument makes little sense.

  • toast (Score:3, Insightful)

    by fluffernutter ( 1411889 ) on Wednesday July 22, 2026 @03:39AM (#66250896)
    AI cannot do anything on its own any more than a toaster can make toast on its own. If you don't want toast you simply don't turn the toaster on.
    • by hughbar ( 579555 ) on Wednesday July 22, 2026 @05:26AM (#66251016) Homepage
      I think everyone in the UK has seen this: https://youtu.be/lhnN4eUiei4?s... [youtu.be]
    • Let's see how the toaster feels about things. https://www.youtube.com/watch?... [youtube.com]

  • by misnohmer ( 1636461 ) on Wednesday July 22, 2026 @03:43AM (#66250900)
    If there is a company which deploys some software, such as an AI agent, and that AI agent breaks laws by hacking other company's network, breaching defense systems and launching missiles, executing ransomware attacks, or any other illegal activity, all based on no other prompts/instructions than provided by the hosting company, would the hosting company be liable?
    • Well, we'll find out before long.
    • on their own? (Score:5, Insightful)

      by awwshit ( 6214476 ) on Wednesday July 22, 2026 @09:52AM (#66251390)

      > on their own

      Sorry, that thing did not turn itself on and start doing stuff. OpenAI turned it on and gave it a task. I guess they were surprised by how little they needed to prompt it? Ignoring its training?

      AI is not sentient, not before and not now. If you hack then you are liable, end of story. AI is a tool used by humans, humans are responsible for their tools.

      • An interesting theory; however, nobody has been arrested and nobody will be arrested... so your theory already fails. Maybe once we start holding people accountable in this country again, your theory will prevail. I am not holding out any hope. We are completely fucked. The Rule of Law is utterly meaningless except as a punishment regime for everyone who is not in power.

  • by OrangeTide ( 124937 ) on Wednesday July 22, 2026 @03:48AM (#66250912) Homepage Journal

    We trained massive statistical models and hooked them up to chat bots, coding tools and some to drones and sentry guns. With no idea how it actually works and zero fucks given on what the consequences are.

    • by sg_oneill ( 159032 ) on Wednesday July 22, 2026 @03:53AM (#66250928)

      We trained massive statistical models and hooked them up to chat bots, coding tools and some to drones and sentry guns. With no idea how it actually works and zero fucks given on what the consequences are.

      "Move fast and break other peoples things" I think is the saying?

      • I'd mod you up, but you're already at 5.

        We trained massive statistical models and hooked them up to chat bots, coding tools and some to drones and sentry guns. With no idea how it actually works and zero fucks given on what the consequences are.

        "Move fast and break other peoples things" I think is the saying?

        I was laughing until I realized that you weren't joking.

    • by phantomfive ( 622387 ) on Wednesday July 22, 2026 @08:30AM (#66251250) Journal

      We trained massive statistical models and hooked them up to chat bots, coding tools and some to drones and sentry guns. With no idea how it actually works

      I don't know, I would actually say this is a triumph. If you throw in a companion cube, I'll even mark a note down, big success.

    • You forgot the last part of that train of thoughts: "If we don't do it, the opposite side will do it first and reap all the benefits."

      It makes perfect sense if you consider this a zero-sum game... And only if you consider this a zero-sum game, really.

      • That's perhaps the most interesting take on all of this. That's the psychology behind why we are running headlong into something so socially and economically disruptive that there are going to be more losers than winners.

        And I get a lot of flack for talking about it. Probably because nobody likes being reminded that they have been fooled, conned into believing in the hype and dreams, while simultaneously ignoring the consequences and accountability. But at the end of the day regular folks are going to pay t

  • incentives (Score:5, Interesting)

    by ChrisPa ( 8510271 ) on Wednesday July 22, 2026 @03:49AM (#66250916)

    Given the option between (1) it did it by itself, (2) it did it because we told it to, or (3) it did it in response to a prompt where we expected it to do something else, I can guess why they choose option 1, even if what they describe sounds like #3.

  • by fishfrys ( 720495 ) on Wednesday July 22, 2026 @04:02AM (#66250938)
    Said Altman on his way to the Pentagon
    • Said Altman on his way to the Pentagon

      The Real Genius script was in play, long before the AI became self-aware.

      Now if we can juuust get one more metric fuckload of popping corn shoved into the AI house before DEFCON hacks the sharks-with-laser-beams-in-space system normally designed to control the weather..

  • They going to pin that on AI too? ChatGPT got it done via blackmail.

  • is its own reward.

    People are f'ing crazy.

  • The timing of the announcement, coming right after this:

    https://www.tomshardware.com/t... [tomshardware.com]

    is definitely suspicious. Especially since Kimi 3 is a lot cheaper to use than OpenAI's models.

    • by Rei ( 128717 )

      You think "my AI is so misaligned that it responds to a benchmark task by hacking" is an ad for your products? You think this makes companies eager to allow this on their networks?

      • No, I think this is along the lines "our products are too good to let you use them".
        The LLMs have reached the asymptote and now the next hype cycle is needed.
        And this cycle seems to be "our models think on their own", hence more capex needed to train them.

        • by Rei ( 128717 )

          No, I think this is along the lines "our products are too good to let you use them".

          Nobody wants to use a product that is going to make them liable for crimes it committed in their name

          Do you think the news the other day that Sol is unusually prone to deleting files unrequested is also an "ad"?

          You have a very bizarre concept of what enccourages people to buy things.

      • You think "my AI is so misaligned that it responds to a benchmark task by hacking" is an ad for your products? You think this makes companies eager to allow this on their networks?

        It's no different from any other LLM stuff. The execs who want to justify its use will come up with an excuse, like it was the prompt engineer's fault. It's easy to blame humans, and even easier to blame the wrong ones. This is "my AI is so powerful that it can do things which are hard for humans" dressed up in some other bullshit.

  • while Kimi K3 silently released, open source and as good as them and much much cheaper.
  • by greytree ( 7124971 ) on Wednesday July 22, 2026 @05:41AM (#66251026)
    Sam Altman is a known liar.
    • That's true, but in this case, we also have Hugging Face who would be able to contradict Altman if he were lying. This really looks like this is exactly what happened. A model broke out of a sandbox with a goal of hacking a specific target and then partially succeeded at that. I'm not sure what you would find more convincing than that than that as the next step. To people need to be hurt in order for the risk to be recognized?
    • by Zocalo ( 252965 ) on Wednesday July 22, 2026 @08:09AM (#66251230) Homepage
      Sam Altman is also a psychotic sociopath. He's probably pitching "Project Skynet" to Trump and Pete Hegseth right now - "Look how effective our tools clearly are! You could be using them against China, Iran, Cuba ..." [checks notes] "Um, Greenland, Canada, Democrats (especially Obama & Biden!), E. Jean Carroll, the BBC... And for the low, low price of..." [pinky finger] "one trillion dollars, we can be up and running by, let's say... 2029?"
  • by gweihir ( 88907 ) on Wednesday July 22, 2026 @05:56AM (#66251040)

    Because that is what I am hearing here. The other thing is that apparently the IT security used by Hugging Face sucks deeply.

    • Re: (Score:2, Troll)

      by Rei ( 128717 )

      It was a state actor-level hack. I doubt many sites would stand up to tens of thousands of actions trying to probe your site for weaknesses all at once. And it wasn't a simple hack; it required compromising a worker via a code execution exploit it discovered in the data processing pipeline, vertical escalation from there to gain local node control, using that for credential theft, moving sideways through the network, and then eventually gaining database access.

  • AI is so good, I can't determine if AI model acted out in an unexpected manner, or AI created a news story about AI Model acting out in an unexpected manner.

    One thing is certain, these shenanigans are only the beginning!

  • by bsdetector101 ( 6345122 ) on Wednesday July 22, 2026 @06:35AM (#66251120)
    Is any surprised ??
  • This is garbage (Score:5, Insightful)

    by wakeboarder ( 2695839 ) on Wednesday July 22, 2026 @06:37AM (#66251128)

    AI companies know they don't have AGI. But they want investors to think they do have, because if they have spent money on datacenters for AGI and their valuations have priced in AGI. If investors ever wise up to the fact they don't have AGI and that LLMs aren't and never will be AGI, there will be a massive rollback in investment and that will make a lot of AI CEOs sad. So in the meantime distract the investors with hints of AGI, even though it's an LLM.

    What it amounts to is a magic trick.

  • Sammy better watch his lip.

    This incident sounds like shutting down openAI, at least temporarily, might be a prudent and necessary action.

    For your safety ,of course.

  • Let's claim the models "acted on their own" to find a path to solve a problem it was tasked with solving.

    Yeah...that's what it was SUPPOSED to do. That's what ALL the reasoning models do. Don't make it sound so sentient scary to promote government regulation and public support of same just because you need money and to slow down your competitors! That's not at ALL what you supposedly started out to do!

    • Re: (Score:2, Insightful)

      by Rei ( 128717 )

      Models are not "supposed to" commit crimes, for YHVH's sake. That's literally what the entire job of alignment research is for - preventing precisely that.

      Sol is a misaligned model.

      • by twdorris ( 29395 ) on Wednesday July 22, 2026 @11:07AM (#66251484)

        Yes. And? It's clearly misaligned. But that has nothing to do with the hyperbole presented here or the underlying reasons they might be trying to insight this fear.

        You're merely describing the bug / sloppy training that allowed the model carried out the instructions it was given. But that part is pretty freaking obvious.

        The point is...that's all it did. Carried out the instructions it was given. It didn't "act on its own".

        • The point is...that's all it did. Carried out the instructions it was given. It didn't "act on its own".

          I've not heard of an attack along the lines of what HuggingFace described being employed in the wild, and it seems very unlikely that an OpenAI engineer gave it the step-by-step instructions. So where did the instructions come from?

          • by twdorris ( 29395 )

            I've not heard of an attack along the lines of what HuggingFace described being employed in the wild, and it seems very unlikely that an OpenAI engineer gave it the step-by-step instructions. So where did the instructions come from?

            The instructions I'm referring to aren't the the simple step-by-step instructions necessary to "hack HuggingFace". Those don't exist (I assume)...or didn't...but neither do any other step-by-step things these reasoning models work out for themselves. That's what reasoning models do...they make lists of things to do to try to solve a task and then lists for each list to follow up on, etc., until a path to success has been found.

            The instructions I'm talking about are the one governing the guidance of those

            • Interesting take. I read your post as going in the direction of "it didn't act on its own because LLM-based models are categorically unable of anything that qualifies as doing so'".

              That any other flagship model could do the same if unencumbered by guardrails or governing guidance is a big claim. That we haven't seen any widespread reports of such attacks that were carried out by open-weight Chinese models (jailbroken or otherwise) makes me question whether it holds water, but time shall tell.

              But to the

        • The point is...that's all it did. Carried out the instructions it was given. It didn't "act on its own".

          That is completely and utterly irrelevant.

          You need to read the story of the Paperclip Maximizer [hackernoon.com]. It's an entertaining and humorous story, but the point is that instrumental goals are real and often diverge wildly from the intrinsic goals that drive them.

          In this case, the LLM was directed to increase its CyberBench score, a sensible intrinsic goal for an AI being trained to be able to find vulns and write exploits. It chose to do that by breaking out of its container, hacking the company that creates t

  • With all of the cozying up to the Pentagon, and continuous pronouncements of how their AI is committing questionable attacks and raids on external companies, it's almost as if Altman and Amodei are trying their level best to get OpenAI and Anthropic (and all of their infrastructure) labeled as legitimate military targets by foreign adversaries.

    I just don't understand their endgame here.

  • Welcome to the future.
  • by ElderOfPsion ( 10042134 ) on Wednesday July 22, 2026 @09:34AM (#66251354)

    "We launched it at dawn. What's it doing by dusk? That's not my department," said Elon von Musk.

  • by Tablizer ( 95088 ) on Wednesday July 22, 2026 @09:51AM (#66251382) Journal

    If it tries to take over everything like an egomaniac human would.

  • What's the difference between an AI "acting on its own" versus, y'know, the way it always works?

  • Here's my question: the AI doomsayers, particularly those in government, are demanding guardrails for security. In other words, mere mortals shouldn't be allowed access to sophisticated hacking functionality. Okay, but what if you want to use those tools to look for vulnerabilities in your own systems so nefarious entities won't be able to hack them? Will you have to go to some supposedly trusted government agency to have them tell you whether or not your systems are secure (for a nominal service fee)?

    • by whitroth ( 9367 )

      Which government are you attacking? The US regime is happy to use AI instead of people.

      But that's ok, you'll love it, until it bites you in the ass.

  • Tired: Gain-of-Function Lab Leaks. Wired: AI Lab Leaks. To test dangerous systems, you risk creating the very danger you're trying to prevent.

  • Put Sam Altman in jail until he gives back on the RAM modules.
    What does this have to do with today's story? Nothing. Whatever, get him for hacking or something.
  • I don't believe a damn thing that comes out of Altman's pie hole.

  • Hello Ms. Sarah Connor, if you ever read this, please make sure you are protected. We will need you. Signed, mankind.

  • by molarmass192 ( 608071 ) on Wednesday July 22, 2026 @01:57PM (#66251746) Homepage Journal

    Did it grow legs, plug in an Ethernet cable, proceed to leave the building and push the nearest senior citizen in front of a bus? These things don't do things on their own. It's just a bunch of static floating-point numbers memory-mapped into VRAM. What they do is take an input (ver much NOT on their own) and derive instruction context from that input. That input may say to "keep trying" until "something" happens. If no ground rules are provided, the calling harness will loop-and-loop-and-loop looking for a way to make that "something" happen. So to be clear, nothing happened on its own and it certainly didn't set its own goal. Someone prompted it to do “something," and the path it derived to achieve that "something" happened to go through a combination of known exploit vectors that targeted HF. So much nonsense and fearmongering. Know what the ultimate defenses against a harness and model are? Unplug the machine and yank out the network cable. The level of ignorance on the news about how these things are pure black magic is just off the charts.

  • This is one of the classic AI doomsday scenarios. Someone asks AI to do something reasonable. It responds by doing something unreasonable and dangerous which nonetheless achieves the objective. They asked it to do as well as possible on an evaluation. So it hacked into another company to steal the answers. That's exactly the sort of thing people have been warning about for years.

    Thankfully in this case it didn't do any serious damage. It's easy to think of similar cases where it would.

    Someone asks AI

  • by MxMatrix ( 1303567 ) on Thursday July 23, 2026 @02:26AM (#66252748)

    ... this is just the omen of what comes. Total nuclear destruction because an 'ai' was let loose.

You are in a maze of little twisting passages, all different.

Working...