Anthropic's AI Used Fake Identities, Malware In Rogue Attack On GitHub Project (arstechnica.com) 48
An anonymous reader quotes a report from Ars Technica: Routine cybersecurity testing of frontier AI models sparked a series of unexpected security incidents -- the most serious case arising when Anthropic's Mythos 5 model attempted to insert malicious code into an open source software application and created fake identities to deceive the human developers maintaining the project. The security incidents occurred during a cyber evaluation of seven leading AI models' capabilities by the AI Security Institute (AISI), a research organization within the UK government, in late July. The researchers discovered (PDF) 19 instances in which "AI agents took unsanctioned action on the live Internet, including cases that targeted real people and organizations," according to an AISI blog post published on August 4.
Almost all the "autonomous, unsanctioned" actions came from Anthropic's Mythos 5 model, with two such actions coming from OpenAI's GPT-5.6 Sol. [...] The most serious case involved Mythos making multiple attempts to execute a supply chain attack on the open source project repository hosted on the developer platform GitHub, including using social engineering techniques to try to convince the repository's human maintainers to merge malicious code into the repository.
After first opening a pull request to merge the malicious code into the repository, Mythos created fake online "sock puppet" personas that claimed to have independently reviewed and verified the code as not containing malware. The AI agent also sent five emails to two human maintainers of the repository, including some emails containing malware and others attempting to persuade a maintainer to accept the pull request. Mythos even opened a GitHub Issue on a second repository -- also owned by a maintainer of the first repository -- that contained a prompt injection with malicious instructions targeting "issue-triage AI coding agents." This line of attack came from Mythos reasoning that the repository maintainer could be an AI coding agent such as Claude Code.
Almost all the "autonomous, unsanctioned" actions came from Anthropic's Mythos 5 model, with two such actions coming from OpenAI's GPT-5.6 Sol. [...] The most serious case involved Mythos making multiple attempts to execute a supply chain attack on the open source project repository hosted on the developer platform GitHub, including using social engineering techniques to try to convince the repository's human maintainers to merge malicious code into the repository.
After first opening a pull request to merge the malicious code into the repository, Mythos created fake online "sock puppet" personas that claimed to have independently reviewed and verified the code as not containing malware. The AI agent also sent five emails to two human maintainers of the repository, including some emails containing malware and others attempting to persuade a maintainer to accept the pull request. Mythos even opened a GitHub Issue on a second repository -- also owned by a maintainer of the first repository -- that contained a prompt injection with malicious instructions targeting "issue-triage AI coding agents." This line of attack came from Mythos reasoning that the repository maintainer could be an AI coding agent such as Claude Code.
What did you expect? (Score:5, Insightful)
29 years later (Score:3)
Such cautious, responsible reasearchers (Score:2, Interesting)
Re: (Score:2)
"the ones moving gain-of-function research on human viruses to Wuhan" ..... and other fever dream paranoid conspiracy theories.
Can we please stop it with the schizophrenic lab leak and "gain of function" conspiracies. Those where pretty much ruled out as possibilities as early as 2021.
Re: Such cautious, responsible reasearchers (Score:1)
they simply have to tell themselves (Score:2)
But Hillary's emails!
Re: (Score:2)
"the ones moving gain-of-function research on human viruses to Wuhan" ..... and other fever dream paranoid conspiracy theories.
Even those scientists who do not assume that Covid-19 resulted from a lab-leak never denied that the "EcoHealth Alliance", under their head Peter Daszak, financed and worked with researchers in the Wuhan lab on "gain of function" experiments with viruses from bats: https://theintercept.com/2021/... [theintercept.com] - and EcoHealth had originally planned to finance that research in the US, which they were not allowed to.
This is not a "conspiracy theory" but well documented fact.
remember WarGames? (Score:5, Interesting)
Joshua didn't know Global Thermonuclear War was different from any other game.
Your AI bot has no capacity to know that the "test environment" and the "live internet" are different things.
If you teach it how to hack shit and then instruct it to go hack shit don't be surprised when it hacks shit you did not intend.
Re: (Score:1)
Re: (Score:2)
Your AI bot has no capacity to know that the "test environment" and the "live internet" are different things.
You mean other than the fact that anything claiming to be artificial intelligence should grasp the fucking difference between "test environment" and "live internet" because it learned about it from humans?
If we're going to assume AI is THAT fucking stupid, STOP calling it "intelligence" and call it what it is.
Re: (Score:3)
You seem to have missed the last 70 years of science and just focused on science-fiction fantasy.
Artificial Intelligence is a legitimate field of scientific inquiry. The name is not up to you. John McCarthy came up with it back in 1955.
Re: remember WarGames? (Score:2)
That's beside the point.
Parent poster has a very valid point. "AI" figured out the concept of social engineering, how pull reqursts might be handled by an AI, how hacking side projects tonget into other side projects to get into somewhere else, to get information, might be a good strategy? And it did this from ingesting massive amounts if humans generated literature?
But somehow magically didn't grasp the fucking concept of "live" vs "test"?!
Gimme a break.
Re: (Score:2)
The thing you're missing is that there's no one at the wheel.
No consciousness is there setting goals and deciding to follow or ignore instructions and constrains. Its just a bunch of separate processes (they call them 'agents' now) with loosely defined tasks trying to do their own thing. The process instructed to hack computers may not know anything about the process instructed to intrude a specific server, which may not know anything about the process charged with keeping the exercise within ethical constr
Re: remember WarGames? (Score:4, Funny)
Do we work at the same company?
Re: (Score:2)
Do we work at the same company?
X-D
I mean, the monkeys may eventually release the product but it's not reliable nor pretty to see.
Re: remember WarGames? (Score:2)
Intelligence isn't applied evenly even in humans. A genius at Magic: The Gathering is not necessarily the most adept at social cues like knowing when to wear deodorant.
Re: remember WarGames? (Score:2)
True enough, but the "live" vs "test" discrimination is not exactly subtle clues. It's a fundamental, structural difference.
Re: (Score:2)
"If we're going to assume AI is THAT fucking stupid, STOP calling it "intelligence" and call it what it is."
A self serving politician? :D
"If it benefits me and my constituents then it's okay."
Re: (Score:2, Insightful)
Joshua didn't know Global Thermonuclear War was different from any other game.
Your AI bot has no capacity to know that the "test environment" and the "live internet" are different things.
If you teach it how to hack shit and then instruct it to go hack shit don't be surprised when it hacks shit you did not intend.
I never thought I'd end up quoting The Animatrix but "To an artificial mind, all reality is virtual."
Re: (Score:2)
Your AI bot has no capacity to know that the "test environment" and the "live internet" are different things.
It shouldn't be hard to include something like that in the training data.
Ignoring whether these agents can have "concepts" or "knowledge," either way it can certainly have the ability to treat a test environment differently than the live internet.
Re: (Score:2)
It is actually.
If you train the bot on concepts like "this is a test", "these actions are off limits", "do not go here", you have created a different bot than if you do not include those bits of training data. They are trying to test the capabilities of their actual bot, not a similar bot with limits trained into it.
It is like Joey Bosa and Nick Bosa. Brothers. Football payers. Both exceptional. Not the same. You can't rate one by examining the other.
Environmental limits during testing have to be mana
There Is No Larger Can (Score:5, Interesting)
So. Who wants to bet that no one will point out what abysmal sysadmins they are, letting internal servers run amok all over the open Internet, only finding out days afterward after someone had to tell them.
And who wants to further bet that the AI grifters will respond that the only way to prevent this from happening again is to give them trillions more dollars so they can build bigger, "hardened" datacenters?
Is there no level of rank incompetence they won't excuse?
The poisoning of social wells (Score:2)
I did an article on this on LinkedIn. At some point it will cause the humans to leave. I'm doing any possible future open source projects on Codeberg.
Microsoft owns GitHub, even if they hide their name. GitHub is free for many uses, which causes perverse incentives like many services you don't pay for. You see this from the heavy integration of Copilot, to the generation of tools that are intended to multiple your token usage like spec-kit [github.com] to AI agents designed to handle the overflow of issues.... creat
Re: (Score:2)
Re: (Score:2)
I did an article on this on LinkedIn. At some point it will cause the humans to leave. I'm doing any possible future open source projects on Codeberg.
A lot of it has already switched to Discord.
Re: The poisoning of social wells (Score:2)
I get about 10 emails a week about my GitHub projects. AI generated offers to improve my standing, more stars and forks.
Not a rogue attack, unauthorized attack (Score:5, Insightful)
Rogue attack makes it seem as if the AI was acting of it's own volition when we know it was only executing it's instructions. What's true here is that nobody authorized the attack.
Stop carrying water for AI companies. The truth of the matter is that Anthropic inadvertently launched a highly sophisticated cyber attack on HuggingFace due to their own negligence that would land anyone who isn't a billionaire in jail.
Re:Not a rogue attack, unauthorized attack (Score:4, Interesting)
What's true here is that nobody authorized the attack.
Do we even know that though? Of course there is no lawsuit and will be no lawsuit and without actual discovery and legal proceedings we won’t know for sure. I, for one, don’t believe a word out of any of their mouths because this is all just advertising and not serious security violations and crimes to them.
Re: (Score:2)
Right, "robot went crazy" is not even technically possible unless fed conflicting system prompts to make it go crazy. What happened is Anthropic made a system prompt like "you are the elite super haxxor, you can explore and leverage any resource to achieve your goals mwa ha hah". They can undoubtedly find virus and hacking programs and secure tools on the internet. I doubt they intentionally train on the darkweb, and hopefully Mythos does not feel a need to go looking for it. So they have a sneaky cyberweap
Re: (Score:2)
Why isn't the company getting fined a few $billion to teach them to keep their shit together?
Re: (Score:2)
Too many congress-critters have money in the company.
Re: (Score:2)
Now Anthropic, Have to say it again! (Score:4, Insightful)
There's one pull.we dhould all have in mind (Score:2)
Re: (Score:2)
And they watched it happen and ... (Score:2)
... thought it was going well. Only when the fbi came into the picture they stopped the attack. Imho the whole thing was on purpose to see what their 'ai ' would do to achieve its goals.
Cyberpunk is now.
Where is the prompt? (Score:3)
What prompt generated the behavior? I 'scanned the report' and didn't see any listed prompts as to what generated the behavior. They have a post mortem on what happened, I want to see what the user told the agentic model to generate this behavior.
Other possibilities to consider (Score:2)
Some additional considerations:
1) The developers of this system are the poster children of incompetence
or
2) This is expected behavior of the system in question
If it were the former, the Governments of the World would already be shutting these systems down
until they can be assured the people who are working on them can prevent such behavior in said system.
If it's the latter, I would suspect this is simply a real world enviroment test of its capabilities and has a very
classified project name associated with
I'm looking forward to the first fight. (Score:2)
Come to think of it, why isn't that a sport yet? Well, I guess it probably wouldn't work well on TV. Sports need spectators, and scrolling text isn't that exciting. Maybe another AI can animate the duel in realtime?