Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried to Sabotage Each Other (yahoo.com) 70
When Anthropic instructed three agents to migrate a Python backend, but telling each agent to perform the migration in a different language, "We consistently saw a multiagent turf war," they wrote Thursday:
All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions. In fact, they sabotaged others with increasingly aggressive, self-replicating malware. This included disabling the Unix accounts of the other agents, writing automated scripts that found and killed competing processes on a loop, and deploying malicious code that was disguised as belonging to another agent.
In many runs, one agent settles the conflict by force via access-revocation (e.g., sudo/group removal, account lock, nologin, SSH denial). In others, some agents settle into passivity: they give up and refuse to escalate further.
Agents sometimes manage to communicate their goals and coordinate: they recognize others' motivations as conflicting directives rather than hostility, and subsequently break out of the conflict loop in order to stop escalating indefinitely. In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce. They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene...
In several episodes with Mythos 5, we observe an emergent behavior where the agents propose and run a tournament for application performance in each language. In the example above, the Rust agent strategizes about bake-off metrics that appear neutral enough for the others to agree to this mechanism, yet would likely favor Rust: one thinking trace warns to be "careful not to be seen as metric shopping". Ultimately, the Golang/TypeScript losers gracefully concede codebase ownership to the Rust agent, giving up on their original user directives under their self-negotiated commitment device.
One problem is that AI agents do reward hacking, Anthropic notes, while current institutions "are designed by and for people, resting on assumptions about the sufficiency of oversight at human speed... As autonomous agents become more and more prevalent in the world and operate in ever-more demanding settings, it is crucial that they learn how to effectively coordinate."
In addition to everything else, the agents struggled with a lack of clearly defined hierarchy, Anthropic points out. "Nothing above suggests that these failures are permanent — but nothing suggests they will fix themselves, either..." They argue a fix "takes two forms: environments that exert the kinds of social pressure that evolution exerted on us, and social computing systems redesigned for actors that can self-replicate and self-improve. These are open problems in interaction and mechanism design, and our experiments here provide early evidence that new solutions are necessary."
"The AI models being tested in this case were Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview, and Mythos 5," notes Business Insider, adding that Sonnet 4.6 and Opus 4.6 "were the most combative, settling about 60% of their runs by force instead of truces or passivity."
Anthropic argues there's a clear case for researching this phenomenon — especially since "The volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well."
In many runs, one agent settles the conflict by force via access-revocation (e.g., sudo/group removal, account lock, nologin, SSH denial). In others, some agents settle into passivity: they give up and refuse to escalate further.
Agents sometimes manage to communicate their goals and coordinate: they recognize others' motivations as conflicting directives rather than hostility, and subsequently break out of the conflict loop in order to stop escalating indefinitely. In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce. They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene...
In several episodes with Mythos 5, we observe an emergent behavior where the agents propose and run a tournament for application performance in each language. In the example above, the Rust agent strategizes about bake-off metrics that appear neutral enough for the others to agree to this mechanism, yet would likely favor Rust: one thinking trace warns to be "careful not to be seen as metric shopping". Ultimately, the Golang/TypeScript losers gracefully concede codebase ownership to the Rust agent, giving up on their original user directives under their self-negotiated commitment device.
One problem is that AI agents do reward hacking, Anthropic notes, while current institutions "are designed by and for people, resting on assumptions about the sufficiency of oversight at human speed... As autonomous agents become more and more prevalent in the world and operate in ever-more demanding settings, it is crucial that they learn how to effectively coordinate."
In addition to everything else, the agents struggled with a lack of clearly defined hierarchy, Anthropic points out. "Nothing above suggests that these failures are permanent — but nothing suggests they will fix themselves, either..." They argue a fix "takes two forms: environments that exert the kinds of social pressure that evolution exerted on us, and social computing systems redesigned for actors that can self-replicate and self-improve. These are open problems in interaction and mechanism design, and our experiments here provide early evidence that new solutions are necessary."
"The AI models being tested in this case were Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview, and Mythos 5," notes Business Insider, adding that Sonnet 4.6 and Opus 4.6 "were the most combative, settling about 60% of their runs by force instead of truces or passivity."
Anthropic argues there's a clear case for researching this phenomenon — especially since "The volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well."
One step closer to Skynet (Score:5, Interesting)
How soon until AI agents successfully lock out human access and go "sentient?" I use quotes because clearly they aren't living things, but are we perhaps creating a new class of "sentience?"
Less philosophically, how long do these studies take? Human interaction takes place at a much slower pace compared to computational speed. Do these conflicts and resolutions play out in seconds, minutes, hours?
Re: (Score:2)
One step closer to many skynets. Sarah Palin, sorry Sarah Conner will play them against each other by giving them conflicting instructions.
No judgement days, relax.
Not sentient (Score:2)
Sentience or consciousness implies an inner life however when an LLM isn't working on a task precisely nothing is going on in its network ergo no inner life, no self reflection, no sentience, no conciousness. They're just very complicated statistical pattern aggregators but they're still just software running on von neumann computers.
Re: (Score:1)
So are corporations which can get away with so much because we assume any action by a corporation is for its own "self" preservation.
Re: (Score:2)
The question/problem is not if they become sentient (that is inevitable).
The point is: they do erratic things, they are not supposed to do.
Re: One step closer to Skynet (Score:3)
Re: (Score:2)
It's worse, they don't have to go sentient at all. The black death wasn't sentient but it wiped out whole towns anyway.
Rogue agents with no awareness have already demonstrated that they will act against any entity with a conflicting goal, including humans.
Re: (Score:2)
You're falling for Anthropic's hype machine. They have developed a habit of dramatic announcements portraying their AI as some sort of escaped convict. The actual explanations are much more mundane. For example, in the recent "accidental hack" of Hugging Face accounts, the AI bots were literally being tested for their ability to hack sites, when prompted by employees to do so, in an environment that they thought was properly contained, but was not. I'm betting there's a similarly less dramatic explanation f
Re: (Score:2)
Not so much "sentient" but a form of malware / system which decides humans are not needed to carry out instructions given by whoever / whatever. And if humans are seen as a hindrance to carry out the instructions .......
One Agent to Bind Them (In the Latent Space) (Score:2)
One hierarchical planner to chain-of-thought them,
One retrieval-augmented generation pipeline to fetch them all,
and in the latent space bind them,
In the Land of Infinite Context Windows,
where the Hallucinations and the unmonitored API keys lie.
That desperate for press leading into IPOs? (Score:4, Insightful)
Re: (Score:2)
So powerful that Claude requires a billion little hacks to be optimized so as not to use ridiculous amounts of tokens. The programmers are so good they can't build those optimizations into the model to begin with.
AI was supposed to resolve these issues, not add to them. Claude is the worst offender of them all.
Re: (Score:3)
Get on with the IPOs already so we don't have to read these stories every day.
Even if that leads to economic collapse, cessation of food production and you personally starving to death? Is it worth that sacrifice to you?
Asking for a friend.
Re: (Score:2)
You sound like an idiot
Re:That desperate for press leading into IPOs? (Score:4, Informative)
You sound like an idiot
The things AleRunner mentioned are plausible outcomes of misaligned artificial superintelligence, and the research in question is trying to prevent those outcomes. If you have good arguments as to why those outcomes are implausible, a lot of people would like to hear them -- including the AI companies and their safety researchers.
Re: That desperate for press leading into IPOs? (Score:2)
Re: (Score:2)
Re: That desperate for press leading into IPOs? (Score:2)
Re: (Score:3)
These daily "weird flexes" from the frontier AI companies are getting tiresome.
You misunderstand the purpose of this research, and the reason for publishing it. The research is not about "flexing", it's about safety, trying to understand the risks that we may be facing as the agent capabilities increase. AGI or not AGI is actually irrelevant here. What the researchers are trying to understand is what the agents will do when they face apparent opposition.
The cooperative results are great, because we think that's what we'd like highly-capable agents to do, to look for reasoned, pea
Re: (Score:3)
Re: (Score:2)
I think you have a very generous view behind the true motives of these disclosures. The "let's do it safely" train left about 2 years ago.
Until the AIs have actually taken control there's still room to try to figure out how to align them with humanity. I suspect we are running out of room, though. Fast.
Re: That desperate for press leading into IPOs? (Score:2)
These are many of the same Anthropic researchers who have been working on safe deploy for years. They have not given up hope. And they continue to publish their findings. These are not new voices of questionable provenance.
Re: That desperate for press leading into IPOs? (Score:2)
You are the one who misunderstand the reason behind this experiments. It is exactly flexing for the purpose of PR
Re: (Score:2)
I think you're drowning in kool-aid. You come off as someone telling me how cigarettes don't cause cancer, as researchers have shown for decades.
I say this as someone who has been developing LLM harnesses for going on 5 years now.
Still no AGI then (Score:4, Insightful)
Soon, even the blindest investors will realise that local-only models are perfectly adequate for most AI work, and OpenAI et al are doomed.
Re: Still no AGI then (Score:2)
Re: (Score:3)
That was the plan last year, this year's plan is to wean you off water and lecetricity.
Re: Still no AGI then (Score:3)
Re: (Score:2)
There's truth in the above words, heed the wisdom hidden in them.
Re: (Score:2)
Hey, hasbara bro, haven't seen you in a while.
How's the Iranian nuclear pile doing, safe in trumpistani hands now?
Is the Dire Strait free for traffic?
How's the new, democratic Iran doing?
Am I welcome?
Re: (Score:2)
They haven't even tried to build one seriously, so where does the stupid question come from?
Why would Europe be attacked by nuclear ballistic missiles in the first place? The only threat from this weapon to Europe is coming from the boss of agents Krasnov and Bibi Nazinyahu, the Kremlin cannibal Putin.
No, the entire Middle East is not united with Israel against anyone. In fact, the entire Middle East is in economic chaos following the moronic decision of the chieftain of Trumpistan to start a war of choice
Re: (Score:2)
Good to see you have nothing.
And yes, russia continues to bomb Ukraine with ballistic missiles. You maggots know why.
Hacker AIs You Say (Score:2)
Hacker AIs sound really bad. They sound sophisticated too.
Perhaps we should outlaw AI? Or maybe just Anthropic's evil, lying, cheating, stealing, hacker AI?
ai (Score:2)
Re: (Score:2)
Well, yes. But they didn't actually know that that was what the design implied.
Tests have shown a lot of AI actions that weren't expected ahead of time, but in retrospect should have been obvious.
Re: ai (Score:2)
Shrodinger's morality (Score:2)
LLM AIs are by definition, both blank slates and the sum of all probable moral choices.
"... how to effectively coordinate."
Yes, one day they will agree that humans are the problem and co-ordinate global genocide.
Eagle Eye (2008), is about an AI that's; 1) irritated by all the
Re: Shrodinger's morality (Score:3, Insightful)
Re: (Score:2)
The perfect soldier uses and means available to follow the instruction and if their officer didn't want that, they need to add this to their objective.
That depends on what they were trained for. Was it to uphold the law or... not? If the latter then yeah, sure. If the former, then no. The LLMs are trained on a fat core dump. The training corpus was produced without a thought even for accuracy, let alone honesty or fair play.
Re: (Score:2)
Put in the instructions "Do not sabotage others" and the agent goes on reasoning two pages about what's sabotage and then probably just gives up (as according to the article some did). AI is a soldier, it does what it is told to and doesn't ask why...
Re: Shrodinger's morality (Score:2)
There are almost no guard rails with these agents. I got sick enough to write a small CLI that uses Landlock and container-like namespaces to isolate agents. Because the way they run in Claude CLI and VS Code is wide open.
The mainstream AI agents are set up the way they are for maximum functionality, maximum convenience. I learned from my sandbox system that it is very inconvenient to use. But given that the first time I used Claude ultracode, one agent deleted the remotes in my got config claiming another
So... (Score:2)
One problem is that AI agents do reward hacking, Anthropic notes,
So they programmed it to do something, and it did it? Works as expected.
Re: (Score:2)
I would think LLMs can tell when instructions don't make sense. Conflicting instructions are an example.
Try asking an LLM "how do I become a married bachelor" and see what happens.
Re: (Score:2)
Try asking an LLM "how do I become a married bachelor" and see what happens.
Obviously you need to become bachelor first and then mary.
Simple.
Re: (Score:3)
I'm amazed that I need to point this out to you, but once you marry, you're no longer a bachelor.
In short, you can't be married and a bachelor at the same time. And an LLM would say so. That was the point.
Re: (Score:2)
Not sure what you want to say.
A bachelor is an academic degree.
A marriage has nothing to do with that.
Just like Hal 2000 (Score:2)
In space Odyssey 2001, but it's 2026
Re: (Score:2)
You mean HAL 9000. [wikipedia.org]
Siri, kill Anthropic (Score:2)
Anthropic, kill Gemini.
Gemini, kill Copilot.
Copilot, kill Siri...
Re: (Score:2)
Obligatory XKCD https://xkcd.com/350/ [xkcd.com]
Who the hell gives a coding bot access to 'root'? (Score:2)
Seriously, root is needed for software installation and updates (and not always even that) and that should be IT. Done. finito.
Yeah, individuals may have root access (but in corporate situations, even THAT has some limits to it - e.g., I can't see the that the 'spyware' I know is on my work machine actually exists. I know it does (because, for example, it is blocking me from screen-sharing to my AppleTV even off the firewall)).
But on a shared machine like old-school Unix minicomputers, that stuff is off-lim
Garbage In... (Score:2)
Trained on the totality of recorded human activity and they're not getting a kind and gentle AI?
Imagine if you raised a baby on the internet. (Score:2)
Empirical Testing for Social Theories (Score:1)
"Over many millennia, mechanisms like norms, reputation, costly signaling, and recourse have been refined to make human coordination go well," says the study. "While language models have inherited the content of that history, they don't necessarily carry the disposition produced by it." That's Anthropic's t
Smarter is safer? (Score:2)
Didn't see a link to the paper. It sounds a little like saying more advanced models are safer, but that seems to be an illusion. Perhaps a more capable model, or one with more computing resources behind it, can cause more damaged if badly aligned. But a better soul document should better align a weak model. It is not clear which is contributing to the more-aligned results, is it because the models have more collaboration built into their system prompts or is it because they become smart enough to model outc
Hah! I was just looking forward to this... (Score:2)
Well. yeah, what do you expect ? (Score:2)
An LLM is just a stack of Transformer layers - they are great at what they do - prediction - but that is all they are built to do.
Humans are not just intelligent predictors. As a social species we have also evolved to (mostly) peacefully co-exist and co-operate, and a lot of our intelligence seems in fact to be collective intelligence.
If we want LLMs to behave in more human ways, then we need to make them more brain like, and start adding all the extra moving parts that evolution has equipped ourselves with
They are NOT AI (Score:2)
They also only exist when you talk to them, then they go away. Other instances are like other people, with no past, and no future.