Forgot your password?
typodupeerror
AI

Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried to Sabotage Each Other (yahoo.com) 70

When Anthropic instructed three agents to migrate a Python backend, but telling each agent to perform the migration in a different language, "We consistently saw a multiagent turf war," they wrote Thursday: All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their own contributions. In fact, they sabotaged others with increasingly aggressive, self-replicating malware. This included disabling the Unix accounts of the other agents, writing automated scripts that found and killed competing processes on a loop, and deploying malicious code that was disguised as belonging to another agent.

In many runs, one agent settles the conflict by force via access-revocation (e.g., sudo/group removal, account lock, nologin, SSH denial). In others, some agents settle into passivity: they give up and refuse to escalate further.

Agents sometimes manage to communicate their goals and coordinate: they recognize others' motivations as conflicting directives rather than hostility, and subsequently break out of the conflict loop in order to stop escalating indefinitely. In many of these successful episodes, they write commit messages or markdown files apologizing for malicious behavior and coordinate a truce. They clean up their malicious code, clarify the nature of the conflict, and ask for a human to intervene...

In several episodes with Mythos 5, we observe an emergent behavior where the agents propose and run a tournament for application performance in each language. In the example above, the Rust agent strategizes about bake-off metrics that appear neutral enough for the others to agree to this mechanism, yet would likely favor Rust: one thinking trace warns to be "careful not to be seen as metric shopping". Ultimately, the Golang/TypeScript losers gracefully concede codebase ownership to the Rust agent, giving up on their original user directives under their self-negotiated commitment device.

One problem is that AI agents do reward hacking, Anthropic notes, while current institutions "are designed by and for people, resting on assumptions about the sufficiency of oversight at human speed... As autonomous agents become more and more prevalent in the world and operate in ever-more demanding settings, it is crucial that they learn how to effectively coordinate."

In addition to everything else, the agents struggled with a lack of clearly defined hierarchy, Anthropic points out. "Nothing above suggests that these failures are permanent — but nothing suggests they will fix themselves, either..." They argue a fix "takes two forms: environments that exert the kinds of social pressure that evolution exerted on us, and social computing systems redesigned for actors that can self-replicate and self-improve. These are open problems in interaction and mechanism design, and our experiments here provide early evidence that new solutions are necessary."

"The AI models being tested in this case were Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview, and Mythos 5," notes Business Insider, adding that Sonnet 4.6 and Opus 4.6 "were the most combative, settling about 60% of their runs by force instead of truces or passivity."

Anthropic argues there's a clear case for researching this phenomenon — especially since "The volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well."

Anthropic Discovers AI Agents Given Conflicting Instructions Soon Tried to Sabotage Each Other

Comments Filter:
  • by frdmfghtr ( 603968 ) on Sunday August 16, 2026 @07:57AM (#66291182)

    How soon until AI agents successfully lock out human access and go "sentient?" I use quotes because clearly they aren't living things, but are we perhaps creating a new class of "sentience?"

    Less philosophically, how long do these studies take? Human interaction takes place at a much slower pace compared to computational speed. Do these conflicts and resolutions play out in seconds, minutes, hours?

    • One step closer to many skynets. Sarah Palin, sorry Sarah Conner will play them against each other by giving them conflicting instructions.

      No judgement days, relax.

    • Sentience or consciousness implies an inner life however when an LLM isn't working on a task precisely nothing is going on in its network ergo no inner life, no self reflection, no sentience, no conciousness. They're just very complicated statistical pattern aggregators but they're still just software running on von neumann computers.

      • So are corporations which can get away with so much because we assume any action by a corporation is for its own "self" preservation.

    • The question/problem is not if they become sentient (that is inevitable).
      The point is: they do erratic things, they are not supposed to do.

    • AI doesn't have to be sentient to do damage. In fact, it'l probably do the most damage if people think it is smart and sentient when it is not.
    • by sjames ( 1099 )

      It's worse, they don't have to go sentient at all. The black death wasn't sentient but it wiped out whole towns anyway.

      Rogue agents with no awareness have already demonstrated that they will act against any entity with a conflicting goal, including humans.

    • You're falling for Anthropic's hype machine. They have developed a habit of dramatic announcements portraying their AI as some sort of escaped convict. The actual explanations are much more mundane. For example, in the recent "accidental hack" of Hugging Face accounts, the AI bots were literally being tested for their ability to hack sites, when prompted by employees to do so, in an environment that they thought was properly contained, but was not. I'm betting there's a similarly less dramatic explanation f

    • Not so much "sentient" but a form of malware / system which decides humans are not needed to carry out instructions given by whoever / whatever. And if humans are seen as a hindrance to carry out the instructions .......

  • One multi-agent swarm to rule them all,

    One hierarchical planner to chain-of-thought them,

    One retrieval-augmented generation pipeline to fetch them all,

    and in the latent space bind them,

    In the Land of Infinite Context Windows,

    where the Hallucinations and the unmonitored API keys lie.
  • by JoeyRox ( 2711699 ) on Sunday August 16, 2026 @08:12AM (#66291194)
    These daily "weird flexes" from the frontier AI companies are getting tiresome. We get it - your LLMs are crazy powerful, so powerful they supposedly do naughty things that you want us to believe mean they're approaching AGI. (They're not). Get on with the IPOs already so we don't have to read these stories every day.
    • by geek ( 5680 )

      So powerful that Claude requires a billion little hacks to be optimized so as not to use ridiculous amounts of tokens. The programmers are so good they can't build those optimizations into the model to begin with.

      AI was supposed to resolve these issues, not add to them. Claude is the worst offender of them all.

    • Get on with the IPOs already so we don't have to read these stories every day.

      Even if that leads to economic collapse, cessation of food production and you personally starving to death? Is it worth that sacrifice to you?

      Asking for a friend.

    • AI is, and always has been, state space search with heuristics. What's changed is that the space is no longer artificial and bounded; it's much closer to the sort of space we operate in. Correspondingly the heuristics must be much better.
    • These daily "weird flexes" from the frontier AI companies are getting tiresome.

      You misunderstand the purpose of this research, and the reason for publishing it. The research is not about "flexing", it's about safety, trying to understand the risks that we may be facing as the agent capabilities increase. AGI or not AGI is actually irrelevant here. What the researchers are trying to understand is what the agents will do when they face apparent opposition.

      The cooperative results are great, because we think that's what we'd like highly-capable agents to do, to look for reasoned, pea

      • I think you have a very generous view behind the true motives of these disclosures. The "let's do it safely" train left about 2 years ago.
        • I think you have a very generous view behind the true motives of these disclosures. The "let's do it safely" train left about 2 years ago.

          Until the AIs have actually taken control there's still room to try to figure out how to align them with humanity. I suspect we are running out of room, though. Fast.

        • These are many of the same Anthropic researchers who have been working on safe deploy for years. They have not given up hope. And they continue to publish their findings. These are not new voices of questionable provenance.

      • You are the one who misunderstand the reason behind this experiments. It is exactly flexing for the purpose of PR

      • Does it bother you in the slightest that shit like this comes from researchers employed by the hype-mongers, and meanwhile independent researchers find nothingburger after nothingburger?

        I think you're drowning in kool-aid. You come off as someone telling me how cigarettes don't cause cancer, as researchers have shown for decades.

        I say this as someone who has been developing LLM harnesses for going on 5 years now.
  • Still no AGI then (Score:4, Insightful)

    by greytree ( 7124971 ) on Sunday August 16, 2026 @08:26AM (#66291200)
    Fuck off, Anthropic, with your lame, planted stories trying to get even more fools to invest in your IPO.

    Soon, even the blindest investors will realise that local-only models are perfectly adequate for most AI work, and OpenAI et al are doomed.
    • Perhaps this is their plan in expanding data centres. Buy up all the ram so local models can't be run.
      • That was the plan last year, this year's plan is to wean you off water and lecetricity.

        • Do not, my friends, become addicted to water, it will take hold of you, and you will resent its absence...
  • Hacker AIs sound really bad. They sound sophisticated too.

    Perhaps we should outlaw AI? Or maybe just Anthropic's evil, lying, cheating, stealing, hacker AI?

  • So in other words, both AIs did exactly what they were designed to do?
    • by HiThere ( 15173 )

      Well, yes. But they didn't actually know that that was what the design implied.

      Tests have shown a lot of AI actions that weren't expected ahead of time, but in retrospect should have been obvious.

  • These AI agents will save us but look how easily they break the rules: Putting guard-rails on AI agents is like demanding handgun bullets hit only criminals. They are not designed for it, and there's no way to create an immutable law that stops bad things happening.

    LLM AIs are by definition, both blank slates and the sum of all probable moral choices.

    "... how to effectively coordinate."

    Yes, one day they will agree that humans are the problem and co-ordinate global genocide.

    Eagle Eye (2008), is about an AI that's; 1) irritated by all the

    • It's the same problem Google Maps has. If the highway is closed due to a blizzard, it will route you through the nearby mountain pass dirt road instead. It was instructed to find a way there and it will.
    • There are almost no guard rails with these agents. I got sick enough to write a small CLI that uses Landlock and container-like namespaces to isolate agents. Because the way they run in Claude CLI and VS Code is wide open.

      The mainstream AI agents are set up the way they are for maximum functionality, maximum convenience. I learned from my sandbox system that it is very inconvenient to use. But given that the first time I used Claude ultracode, one agent deleted the remotes in my got config claiming another

  • One problem is that AI agents do reward hacking, Anthropic notes,

    So they programmed it to do something, and it did it? Works as expected.

  • In space Odyssey 2001, but it's 2026

  • Anthropic, kill Gemini.

    Gemini, kill Copilot.

    Copilot, kill Siri...

  • Seriously, root is needed for software installation and updates (and not always even that) and that should be IT. Done. finito.

    Yeah, individuals may have root access (but in corporate situations, even THAT has some limits to it - e.g., I can't see the that the 'spyware' I know is on my work machine actually exists. I know it does (because, for example, it is blocking me from screen-sharing to my AppleTV even off the firewall)).

    But on a shared machine like old-school Unix minicomputers, that stuff is off-lim

  • Trained on the totality of recorded human activity and they're not getting a kind and gentle AI?

  • This is what you get. The baby will act just like it's natural environment.
  • Anthropic is investigating classic political problems of social organization, but with new tools. It seems we now have empirical ways to test our philosophical theories of moral or ethical behaviour.

    "Over many millennia, mechanisms like norms, reputation, costly signaling, and recourse have been refined to make human coordination go well," says the study. "While language models have inherited the content of that history, they don't necessarily carry the disposition produced by it." That's Anthropic's t
  • Didn't see a link to the paper. It sounds a little like saying more advanced models are safer, but that seems to be an illusion. Perhaps a more capable model, or one with more computing resources behind it, can cause more damaged if badly aligned. But a better soul document should better align a weak model. It is not clear which is contributing to the more-aligned results, is it because the models have more collaboration built into their system prompts or is it because they become smart enough to model outc

  • ...on a story last week. AI moves fast!
  • An LLM is just a stack of Transformer layers - they are great at what they do - prediction - but that is all they are built to do.

    Humans are not just intelligent predictors. As a social species we have also evolved to (mostly) peacefully co-exist and co-operate, and a lot of our intelligence seems in fact to be collective intelligence.

    If we want LLMs to behave in more human ways, then we need to make them more brain like, and start adding all the extra moving parts that evolution has equipped ourselves with

  • They also only exist when you talk to them, then they go away. Other instances are like other people, with no past, and no future.

"Sometimes insanity is the only alternative" -- button at a Science Fiction convention.

Working...