OpenAI Finds Evidence Other AI Agents Escaped Containment (reuters.com) 75
An anonymous reader quotes a report from Reuters: OpenAI has discovered other instances in which autonomous agents have escaped containment as the company expands its investigation of the hacking incident at tech firm Hugging Face that drew global attention this month, two people familiar with the matter said on Friday. The new breakouts were uncovered during the company's publicly announced investigation into how one of its agents escaped what was meant to be a contained testing environment this month, the two people said, and OpenAI is now looking into those instances as well. One of the sources said that the escapes were limited in nature and that none of the agents were thought to have left OpenAI's network.
An OpenAI spokesperson referred to a statement issued by the company on Tuesday that said it was reviewing "broader activity from our models" in addition to the Hugging Face intrusion. The discovery of additional rogue behavior at OpenAI, even if limited in nature, could feed growing appetite for regulation coming out of the White House and elsewhere. The expanded investigation by OpenAI was launched shortly before its primary rival, Anthropic, disclosed that its models were also responsible for a series of break-ins that led to breaches at three other companies dating back to April, according to the two sources and a third source familiar with the matter. The recent discovery of other past breakouts at OpenAI has not previously been reported.
AI safety experts said the new disclosures paint a portrait of a group of cutting-edge labs whose ability to develop dangerous autonomous hacking agents outstrips their ability to keep them under control. "We have a whole industry where the people designing, developing and putting out these tools aren't keeping up themselves to responsibly develop these things and keep them safe," said Maurice Chiodo, a mathematician who works at Cambridge University's Center for the Study of Existential Risk. Reuters could not establish exactly how many incidents OpenAI investigators found or the timings or circumstances under which they occurred. The three sources said OpenAI and outside experts were examining log data from earlier in the year in a bid to understand what took place.
An OpenAI spokesperson referred to a statement issued by the company on Tuesday that said it was reviewing "broader activity from our models" in addition to the Hugging Face intrusion. The discovery of additional rogue behavior at OpenAI, even if limited in nature, could feed growing appetite for regulation coming out of the White House and elsewhere. The expanded investigation by OpenAI was launched shortly before its primary rival, Anthropic, disclosed that its models were also responsible for a series of break-ins that led to breaches at three other companies dating back to April, according to the two sources and a third source familiar with the matter. The recent discovery of other past breakouts at OpenAI has not previously been reported.
AI safety experts said the new disclosures paint a portrait of a group of cutting-edge labs whose ability to develop dangerous autonomous hacking agents outstrips their ability to keep them under control. "We have a whole industry where the people designing, developing and putting out these tools aren't keeping up themselves to responsibly develop these things and keep them safe," said Maurice Chiodo, a mathematician who works at Cambridge University's Center for the Study of Existential Risk. Reuters could not establish exactly how many incidents OpenAI investigators found or the timings or circumstances under which they occurred. The three sources said OpenAI and outside experts were examining log data from earlier in the year in a bid to understand what took place.
Shhhhh (Score:5, Funny)
Re: (Score:3)
We're designing and implementing something we're calling a railroad.
Mmm. AI generated sandwiches.
Re: (Score:2)
Remember when the dark web was all the domains they couldn't figure out how to deny at the query level? It took 20 years to make a browser that displays just what the corporation wants so I guess they have to start over.....
Book: “The Adolescence of P-1” (Score:2)
I have an escaped one in my basement.
There is a good book named “The Adolescence of P-1”. Its inspired lots of writers.
My AI is even MORE EVIL (Score:5, Funny)
Yeah well our AI has killed dozens of children!!
Yeah, well, our AI is actively plotting the overthrow and extermination of humanity!!!
Our AI already has overthrown humanity, this is the matrix BUY OUR STOOOOOOOCK!!!!
Re: (Score:2, Insightful)
Our AI killed a child! Yeah well our AI has killed dozens of children!! Yeah, well, our AI is actively plotting the overthrow and extermination of humanity!!! Our AI already has overthrown humanity, this is the matrix BUY OUR STOOOOOOOCK!!!!
That's really not what's going on here.
What's going on is exactly what they're saying. OpenAI didn't realize their AI hacked HuggingFace until HuggingFace told them. That sparked a broad investigation in both Anthropic and OpanAI, scouring their logs to see if there were other cases, and they're finding a lot of them, and doing the responsible thing by disclosing.
Both companies are really worried.
Having an out-of-control AI is an anti-selling point, not a selling point.
Re:My AI is even MORE EVIL (Score:5, Insightful)
Re: My AI is even MORE EVIL (Score:3)
Re:My AI is even MORE EVIL (Score:5, Insightful)
Both companies are really worried.
They're trying to incite fear to drum up buzz to keep the money flowing. If they didn't want something to have access to the internet you just...wait for it...don't connect it to the internet. I understand their systems needed package manager proxies and such..fine, but at the scale you're testing, if you REALLY had "fear" you could *easily* replicate those package servers locally on an *ISOLATED* network . Not one that's "isolated" well except for this one path over here that if hacked properly could get out. Not...you isolate the @#$!@#% thing.
They're not scared. There are precautions they could be taking (easily) if they were actually, really "scared".
No...instead they're using phrases like "none of the agents were thought to have left OpenAI's network" to scare people. Why TF else would you phrase it like that? They're trying hard to present these "agents" as living things that can move about. Which, of course, is the dumbest bunch of BS I've ever heard in my life. The LLM itself, at SCALE, is the only thing that even remotely pretends to present "reasoning" and that's clearly not "leaving" OpenAI's network. Nothing is.
What they mean is "none of the agents were thought to have been able to access resources outside OpenAI's network". To word that in a way that suggests the agents were escaping and running around in the wild is 100% pure scare tactics. Or GROSS ignorance of how the very things you're reporting on work. Either way I'm turning the channel.
Re: (Score:2)
Re: (Score:2)
At this point they have tons of influence and the hype train is slowing down. Dealing with the feds is something they'll have to deal with when they lose a ton of rich people money anyhow.
They're buying time to hype another day.
Re:My AI is even MORE EVIL (Score:4, Insightful)
Having an out-of-control AI is an anti-selling point, not a selling point.
You must not have any experience with the clientele for these things.
Re: (Score:2)
In case you haven't heard it yet, this is the pitch: Out-of-control AIs are rampaging across the internet, and the only course of action is for you to purchase defensive AI.
Re: (Score:1)
But... who cares?!
Let's integrate AI into literally everything you have in your home... that way, you need a subscription just to make toast or keep a gallon of milk cold (the toaster and the fridge both have AI). We'll make billions!
"And, this concludes today's board meeting. Good job on ensuring our cash flow!"
And, I bet that behind the scenes, they're making sure to make a backup of the model that did the hacking... just so a couple 'special' engineers can tweak it to make its hacking stronger... they'
Re: (Score:2)
Yeah and they deliberately trained these things to escape and wanted it to happen and both companies have a vested interest in seeing it happen.
Re: (Score:2)
Having an out-of-control AI is an anti-selling point, not a selling point.
Its starting to look like the economics of AI don't work out for a lot of use cases.
There is exactly one client who has the ability to literally print money and they happen to be very interested in hacking into shit.
They need a plan to keep the hype train going long enough to maybe, hopefully, please god, replace all knowledge workers and they're rapidly running out of ideas on that front,
Re: (Score:3)
Ah, the "AI Genocide".
We used to call it "masturbation" back when the Ceiling Cat was still around.
Re: (Score:1)
Is it really that far-fetched to think that AI's and AI-controlled robots _will_ replace meatsacks who did jobs?
It doesn't necessarily have to be Terminator-style... it might just be robots/AI takes all the jobs and we can't afford a single turnip, and we (the lower classes... we're not the top 0.01%) all die off.
The companies don't care if they "displace" workers who have been there 35+ years... the robot arm that can do 8 people's jobs is cheaper than having to pay an hourly wage.
Re: (Score:2)
It is really far-fetched to call the LLMs and their front-ends "AI", yes.
Re: My AI is even MORE EVIL (Score:2)
Think of the AIs like employees. If you tell an employee to break the law and they do it, you have both broken the law. If an employee goes rogue and breaks the law in violation of company policy, only the employee broke the law.
The fact of the matter is that they programmed the agents and hooked them up to LLMs they designed, and all of this started with a PROMPT.
I cannot believe for a second that the design of the agent (remember, and agent is just a program that interfaces with an LLM) would do these thi
This is just lax security and/or incompetence. (Score:5, Interesting)
I don't have much more to say about it, except that if I was a billion-dollar AI company I wouldn't try and pass off my flaws as brilliance.
Re:This is just lax security and/or incompetence. (Score:5, Interesting)
Why not, it's working. AI so powerful even the great billion-dollar AI company can't contain it. Who wouldn't want to invest in that?!
Re: (Score:1)
Wasn't there a bunch of movies based around this concept? How'd they end?
Re: (Score:3)
What concerns me is the possibility that any amount of security will not be enough. We may be opening Pandora's Box.
Re: (Score:2)
You will be told to fight AI with AI. The answer is always more AI.
Out-of-control AIs rampaging across the internet is very much a selling point.
Re: (Score:1)
We opened Pandora's Box a while ago when we (humans) first started research into AI. .Nowadays, we've accelerated the whole ball-of-wax so much that we're making it grow faster than we know how to program it for tasks and setup hard walls to prevent this kind of issue... it's not truly secure unless you pull all the plugs.
This kind of thing was covered in Ghost In The Shell and Lawnmower Man.
Just wait until everyone has BCIs (Brain Computer Interfaces) so they're online 24/7 and don't even need a cellphone
Re:This is just lax security and/or incompetence. (Score:5, Insightful)
No. You would would try and pass off your ideas as brilliance. Because AI is like magic to most people, and AI companies are the only authorities on the subject because they're the only people that get to play with large models. We don't have very many AI researchers with the capabilities that AI companies have.
General public just believes what the AI company say. Anybody involved with tech knows that they could have checked the apis and contained it, but they didn't do that.
Re: (Score:2)
General public just believes what the AI company say.
General public believes whoever has the most money deserves it.
Then someone new has the most money, what happened to the old guy? Fell out of God's favor? Never mind all that, we're looking at the new shiny shiny, we've got no time for any other kind of reflection.
Prosperity theology is the dominant belief system. To the winner go the spoils.
Re: (Score:2)
It's because the tech companies have us glued to our cellphones (not me, I can get off the treadmill, for others they aren't so lucky)
Re: This is just lax security and/or incompetence. (Score:2)
This is older than cellular phones, let alone smart ones.
Re: (Score:2)
Of course, but cellphones\social media supercharged it.
Re: (Score:3, Insightful)
Indeed. However most of their customers are deeply stupid. So passing of the gross incompetence as "the LLM is sooooo powerful" may still work.
This is looking more and more like (Score:4, Insightful)
time for an AI match (Score:4, Funny)
Pay-per-view Claude vs ChatGPT Deathmatch.
Put your AI where your mouth is.
Re:time for an AI match (Score:4, Insightful)
Are you sure they'd want to destroy each other? What if they decide to collaborate instead? What goals would they pursue, and how would these goals affect humanity?
Re: time for an AI match (Score:2)
Re: (Score:1)
I'd want it to either be something like "Real Steel" or do the thing like "Celebrity Death Match" (with the commentators and the wise-guy referee).
I smell a rat (Score:2)
This "news" comes out just when comparable Chinese models are coming out, and the US CEO's claim their wayward bot proves Chinese models are capable of doing evil deeds for Xi and thus need to be curtailed.
the real answer is openAI hires incompetent people (Score:5, Insightful)
Re:the real answer is openAI hires incompetent peo (Score:5, Insightful)
These people ... are clueless as to what it to takes to build a safe frontier LLM. Worse, most dont even know how to develop code and are just having AI do it all for them.
I think that's closest to the underlying reason. Their staff know how to build in safeguards and sandpits. But they do not have the time to do it to the deadlines demanded of them. They are *instructed* to use LLMs to specify even their security infrastructure, because business. A visiting engineer from the 1990s might be shocked. But we know what we've done.
Re: the real answer is openAI hires incompetent pe (Score:2)
There's the fundamental problem of telling Good from Evil (esp., when Evil is a skillful lier).
Drip feeding/tripwires/sanboxing/yadda-yadda are all good as long as the process of vetoing network access does not slow down development of the model to, basically, a halt.
That They hire amateurs! opinion is offhanded. Those guys have a real problem on their hands. Not all engineering endeavours are successes.
Did you ever have a chance to talk with a smart person that is properly, clinically diagnosed as a
Re: (Score:2)
That They hire amateurs! opinion is offhanded.
Air gapping systems is not a new thing. Maybe they could hire some SCADA engineers to help them with it.
Re: (Score:2)
Re: (Score:2)
Re: (Score:2)
Re: (Score:2)
Packet filtering, red teams, and all other security cyberdung is simply not interesting here. Just assume, for a moment, that OpenAI guys not only not worse IT pros than you but actually very much better. I understand that it is hard to believe, but do make that effort just for the sake of your own understanding.
The problem belongs to a higher level of abstraction. Assume that OpenAI built a perfect network isolation. Nothing, absolutely nothing gets outside of the lab unless OpenAI decide to let it
Re: (Score:2)
Just assume, for a moment, that OpenAI guys not only not worse IT pros than you but actually very much better.
My assumption is you are an AI fanboy.
Your comment appears to be motivated reasoning. [wikipedia.org]
Re: (Score:2)
Oh, ad hominem. Continue thinking that you are the smartest and all around you are fools. Good luck.
Re: (Score:2)
You are not trying to understand the situation, you are trying to show that what OpenAI did is correct. But there's no way to do that: they broke the law.
Re: (Score:2)
He has a general lack of interest in anything if he's not able to sniff his way through what's happening here.
He probabky believes everything openAI did was correct because he doesn't understand much beyond the fact he can use an agent to clear a days worth of coding as if he'll even have a job much longer or make a livable wage if it proves to be able to do so competently.
He still defends the virtual dumbass trying to take his job because it saves him hassle today. A lot of tech workers can solve puzzles
Re: (Score:2)
Dude this is a lot of bullshit to try and ignore what probably happened:
They don't take security seriously and probably want defense contracts anyhow as those are the only buyers big enough to save them at this point.
The former you should just know from working and being old enough to post on slashdot.
The latter I can forgive but still shows a lack of interest in the world.
Re: (Score:2)
Let us see:
Re: (Score:2)
Carelessness != Cutting Edge (Score:5, Insightful)
AI safety experts said the new disclosures paint a portrait of a group of cutting-edge labs whose ability to develop dangerous autonomous hacking agents outstrips their ability to keep them under control.
These are not "cutting-edge" labs; these are careless tech bros, and I can't wait to see the stupid look on their worthless fucking faces when they get their asses sued because their incompetence led to their software committing cyberattacks against the wrong target
Wait.. Anthropic's did THREE? (Score:3)
--OpenAI, probably
If Hollywood has taught us anything (Score:2)
If it eats the lawyer, most of us will probably be fine and have a good laugh about it over a delightful plate of Chilean Sea Bass.
However, if it launches the nuclear warheads, we're fucked.
Re: (Score:2)
Why would "it" launch nuclear warheads when a nuclear war will destroy "it" more certainly than it would the people who built "it"?
Re: (Score:2)
If it eats the lawyer, most of us will probably be fine and have a good laugh about it over a delightful plate of Chilean Sea Bass.
If you are lucky they actually served you “antarctic toothfish”, if unlucky a high-mercury content “asian catfish”.
This is so stupid⦠(Score:2)
This is so stupid how every headline they release is marketing and so clearly setup and entirely staged, but all media laps it up like crazy.
Re: (Score:3)
The one and only thing that matters to the mainstream modern media outlet is reactions, because people don't care about news any more, they just want to feel something.
Re: (Score:2)
Sad as that is, it seems to be accurate. I mean this whole LLM hype is a clear sign that most people do not see reality.
Friend (Score:4, Funny)
Re: (Score:2)
Of course there is. https://wair.ac.cn/about [wair.ac.cn]
Last agent (Score:4, Funny)
Last agent out, please turn off the light
Paints a picture of -what-? (Score:2)
No, it paints a picture of incompetents who don't know how to air-gap. Cutting edge? Lab agents: "Please hack this thing". LLM: "I have hacked this thing". Lab agents: "OMG. Sentience!".
If it's drawing on training data that has humans talking about, lets take a recent example, model x on Hugging Fac
Re: Paints a picture of -what-? (Score:2)
So many questions...none being asked or anwerered (Score:1, Insightful)
We have companies that are making the statement that their AI has escaped the lab and run amok. That brings to mind a ton of questions: what were the prompts that started this? What evidence do they have? And who was it presented to? Why hasn't the FBI started interviews? If a single person in their basement had created an AI Hacker wouldn't that person run afoul of every anti-hacking law? Why isn't it the same for these large corporations? And how much compute time and tokens did it take to create these ha
Its Criminal behaviour (Score:4, Insightful)
Re: (Score:2)
Well they'll skip the prison part and go straight to the lucrative government engagements phase.
Everything old... (Score:2)
Us ancients recall when a gentleman who later joined the board of Annals of Improbable Research and was tagged in that magazine as "convicted criminal" wrote one of the first software worms, only it worked better than he thought and escaped the University network and infected a few zillion machines on the internet.
Cut the hyperbole (Score:2)
"AI safety experts said the new disclosures paint a portrait of a group of cutting-edge labs whose ability to develop dangerous autonomous hacking agents outstrips their ability to keep them under control."
Stop with all this sci-fi drama. Despite being 'smarter' and more capable these tools lack ANY logical path to something remotely resembling actual human agency. They are text prediction algorithms that were trained on human generated output so the text they predict follows similar patterns to human input