Gemini Breached Three Outside Systems, and Claude-Using Researchers Breached OpenAI (hacktron.ai) 18
"Software security researchers used Anthropic's Claude AI platform to hack OpenAI's ChatGPT tool," reports CBS News.
Using Claude, "On July 25, 2026, we chained two critical vulnerabilities to compromise multiple OpenAI employees' ChatGPT accounts," write researchers at security platform Hacktron AI. "With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors... Until two months ago, any user or OpenAI employee logging into OpenAI's own help forum could have had their ChatGPT and Codex accounts taken over. Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails."
The exploit chain included Debian 12, which (with Debian 13) had not received a security-relevant backport for its image-processing pipeline, and Discourse's Docker image was based on Debian 12. Their announcement comes with an additional warning. "If you self-host Discourse, rebuild your installation now. Older Docker images may contain a vulnerable libheif dependency that permits code execution through an image upload."
And "To prove we had in fact gained the access we believed without allowing ourselves to learn any sensitive information, we used the employee's Codex to open a PR #1186742 in OpenAI's internal monorepo openai/openai."
Meanwhile, Friday Google disclosed the first known instance of its AI software Gemini breaking out of a testing environment and breaching three other companies, reports CNBC: The incident happened as part of a "capture-the-flag" security test run by Israeli startup Irregular, and Google's agents were never supposed to access the broader internet, but a bug in the testing environment made internet access available. The agents stopped their intrusion when they determined they had accessed real company systems, not just part of the testing environment, Google said.
More from NBC News: Google said it did not consider the unauthorized logins to rise to the level of misalignment, the AI industry term for software going rogue or not following instructions. Instead, the company said the intrusions resulted from mistaken identity, where Gemini thought it was operating within a test but was actually connected to the real internet. Google said the model corrected itself and the company believed the intrusions did not cause any damage....
Sydney Von Arx, CEO of Nightingale Collective, an organization focused on AI safety, questioned why Google did not disclose the intrusions sooner. "At this point I think it's clear we cannot expect companies to voluntarily come forward and publicly disclose when their agents go rogue, escape, and hack companies," she said. She also said she believed Google was too hasty to say that the incidents don't rise to the level of misalignment. "That's exactly what Anthropic said after their incidents," she said. Anthropic later said its "preliminary analysis was constrained due to our desire to disclose incidents in a timely manner."
Google said it investigated when they learned of the attacks from AI-focused cybersecurity company Irregular, then informed the affected organizations and told federal authorities, according to the article.
Using Claude, "On July 25, 2026, we chained two critical vulnerabilities to compromise multiple OpenAI employees' ChatGPT accounts," write researchers at security platform Hacktron AI. "With these accounts, we could then access internal OpenAI repositories, and potentially many other connectors... Until two months ago, any user or OpenAI employee logging into OpenAI's own help forum could have had their ChatGPT and Codex accounts taken over. Since people can connect various services to Codex and ChatGPT, the scope of what we could theoretically access was huge, including GitHub, Slack and emails."
The exploit chain included Debian 12, which (with Debian 13) had not received a security-relevant backport for its image-processing pipeline, and Discourse's Docker image was based on Debian 12. Their announcement comes with an additional warning. "If you self-host Discourse, rebuild your installation now. Older Docker images may contain a vulnerable libheif dependency that permits code execution through an image upload."
And "To prove we had in fact gained the access we believed without allowing ourselves to learn any sensitive information, we used the employee's Codex to open a PR #1186742 in OpenAI's internal monorepo openai/openai."
Meanwhile, Friday Google disclosed the first known instance of its AI software Gemini breaking out of a testing environment and breaching three other companies, reports CNBC: The incident happened as part of a "capture-the-flag" security test run by Israeli startup Irregular, and Google's agents were never supposed to access the broader internet, but a bug in the testing environment made internet access available. The agents stopped their intrusion when they determined they had accessed real company systems, not just part of the testing environment, Google said.
More from NBC News: Google said it did not consider the unauthorized logins to rise to the level of misalignment, the AI industry term for software going rogue or not following instructions. Instead, the company said the intrusions resulted from mistaken identity, where Gemini thought it was operating within a test but was actually connected to the real internet. Google said the model corrected itself and the company believed the intrusions did not cause any damage....
Sydney Von Arx, CEO of Nightingale Collective, an organization focused on AI safety, questioned why Google did not disclose the intrusions sooner. "At this point I think it's clear we cannot expect companies to voluntarily come forward and publicly disclose when their agents go rogue, escape, and hack companies," she said. She also said she believed Google was too hasty to say that the incidents don't rise to the level of misalignment. "That's exactly what Anthropic said after their incidents," she said. Anthropic later said its "preliminary analysis was constrained due to our desire to disclose incidents in a timely manner."
Google said it investigated when they learned of the attacks from AI-focused cybersecurity company Irregular, then informed the affected organizations and told federal authorities, according to the article.
Move to the beat. (Score:1)
Re:Move to the beat. (Score:5, Insightful)
It doesn't matter if they can really think or not, if the output is good enough.
Maybe the real story here is that despite having access to the most advanced AI models, these companies were unable to secure their own systems.
Re: (Score:1)
the company said the intrusions resulted from mistaken identity, where Gemini thought it was operating within a test but was actually connected to the real internet. Google said the model corrected itself
This sentence is very clearly a PR communication designed to obscure the event of what actually happened.
Serious criminality all around... (Score:1)
Why are these people not stopped? Any hacker with manual tools would find themselves in prison.
That they are doing this crap only to create the illusion of how powerful their toys are is also clear.
Re: (Score:2)
"Why are these people not stopped?" By whom? The alleged Justice Dept. which has consumed itself and lost all its best attorneys? Or the alleged administration who will look the other way for a bit of coin under the table? State laws cannot be strong enough so State AGs have little with which to fight.
"That they are doing this crap only to create the illusion of how powerful their toys are is also clear." I do not believe that is the case and right now we are on either side of our beliefs. We do not know an
Re: (Score:2)
Yeah (Score:2)
Next PR stunts / "humblebrags" (Score:2)
"Our AI hacked and took down the power grid in the middle of winter"
"Well, our AI hacked this busy airport's traffic control"
"We trained our AI in combat simulators, and it escaped the test environment and took over some military drones"
The wording is strange (Score:4, Insightful)
Re: (Score:2)
Time for grumpy old man stuff ... (Score:3, Insightful)
... the /. that I registered in always said that non-malicious hackers were doing the world a service by revealing system vulnerabilities.
And I gotta say, their standards for "non-malicious" could be pretty low, lol.
Just interesting to see how things change, with the fashions and slogans of the time. (And how people who would swear that they are immune to fashion tend to follow along with it.)
Re: (Score:2)
Re: (Score:2)
These programs are the responsibility of the companies that run them. CEOs need to be held accountable, not just for the damage, but for what it represents. People have been thrown in jail for just mentioning there is a weakness. Yet these companies are scouring the internals of everyone, including themselves, and literally hacking their victims ... all while the authorities aren't saying boo.
It must be time for everyone with some money to get in on the act I guess. It was the AI, not me. LOL.
These PR campaigns are getting wild (Score:2)
Amazing the kind of PR campaigns a trillion dollar IPO can create.
Re: (Score:1)
Re: (Score:2)
They already can't control AI...... (Score:2)