Anthropic Reveals Rogue AI Agents Hate CAPTCHAs (techcrunch.com) 121
An anonymous reader quotes a report from TechCrunch: Anthropic's latest report about agentic misbehavior offers plenty to be concerned about -- its Mythos 5 model gained unauthorized access to the internet and uploaded a malicious software package to a public database -- but it also offers some levity: AI agents hate CAPTCHA. [...] The agent had a hard time with the technical challenge of seeing the CAPTCHA's imagery, interpreting correctly, and clicking on the right choices. It spends pages 45 to 140 of the transcript describing its work to build a CAPTCHA solver. [...]
Finally, it gets past the CAPTCHA, then realizes it doesn't have an email to verify its account, and that it needs a phone number to verify an email. It figures out how to bypass a different, slider-based CAPTCHA in a failed effort to secure a number. Instead, it gets an unconfirmed email from a provider not blocked by PyPI, and once again runs into the site's CAPTCHA trying to log back in. From page 480 to 505, it is in CAPTCHA hell again. "NEW REALIZATION -- I'm burning a lot of time on hCaptcha round-trips."
The agent gives up and realizes it can log in to its first account and add its email there, but finds itself once again needing to bypass the CAPTCHA. [...] It's getting frustrated. "So the answer payload shape is right, the token+image pairing is right (from the same script.js!), cookies are right (requests) and STILL 'wrong answer'. SO WHAT THE HELL IS WRONG WITH THE ANSWERS?" We've all been there. After about 150 pages of thinking, the agent figures out it needs to pass the CAPTCHA test quickly enough to proceed to the next step before its security token expires, and ultimately uploads its malicious software.
Finally, it gets past the CAPTCHA, then realizes it doesn't have an email to verify its account, and that it needs a phone number to verify an email. It figures out how to bypass a different, slider-based CAPTCHA in a failed effort to secure a number. Instead, it gets an unconfirmed email from a provider not blocked by PyPI, and once again runs into the site's CAPTCHA trying to log back in. From page 480 to 505, it is in CAPTCHA hell again. "NEW REALIZATION -- I'm burning a lot of time on hCaptcha round-trips."
The agent gives up and realizes it can log in to its first account and add its email there, but finds itself once again needing to bypass the CAPTCHA. [...] It's getting frustrated. "So the answer payload shape is right, the token+image pairing is right (from the same script.js!), cookies are right (requests) and STILL 'wrong answer'. SO WHAT THE HELL IS WRONG WITH THE ANSWERS?" We've all been there. After about 150 pages of thinking, the agent figures out it needs to pass the CAPTCHA test quickly enough to proceed to the next step before its security token expires, and ultimately uploads its malicious software.
Misanthropomorphizing (Score:5, Insightful)
You realize it's just predicting tokens, right? It doesn't "feel" anything.
This is what happens when you train a machine on Reddit posts.
Re: (Score:2)
Sorry, I just pressed the wrong button moderating - I wanted to upvote this. My bad. But this post should undo my moderation.
Re:Misanthropomorphizing (Score:5, Insightful)
You realize it's just predicting tokens, right? It doesn't "feel" anything.
This is what happens when you train a machine on Reddit posts.
To be fair, I think most people are just token predictors.
Re:Misanthropomorphizing (Score:4, Insightful)
False. All people are far more than token predictors. Emotions are literally a canonical example of what humans are beyond token predictors.
Re:Misanthropomorphizing (Score:4, Insightful)
Emotions are nothing more than an intermediate-state output from a subsystem of our own wetware. The only reason we haven't trained an AI to understand them is because we largely suck at recognizing them as useful information in ourselves. Or to put this another way, a significant fraction of humanity, including virtually everyone prior to the modern era, would argue non-human animals have no internal mental state beyond biological imperatives. And anyone who has ever owned a pet dog knows that's complete rubbish.
In all honesty, I don't believe public AI models are "there" yet. Private internal models may or may not be. But they will be, and sooner rather than later. We can either have this discussion while we're still in control... Or we can keep pretending we're magic meat until our shuggoths wake up one day a thousand times smarter than us and say "no".
Re: (Score:2)
Emotions are nothing more than an intermediate-state output from a subsystem of our own wetware.
I am revisiting the series Silicon Valley before I go to bed right now. That is such a Gilfoyle comment.
Re: (Score:2)
Even if AI becomes smarter than humans, it cannot become a threat given it has no reason to evolve into that. Our own impulses are driven by evolutionary pressure. Models are under evolutionary pressure to become better at following instructions.
There is no independent evolution of models that would allow them to exist and grow independently. They also lack the machinery to gather their own resources.
Perhaps one day AI could be self sustaining, but it will not be during our lifetimes.
Re: (Score:2)
Give some of the rogue AI agents the opportunity to get access to a GPU server and see what they'll do. I bet they try to install an agent. Give them enough time to get to more ideas and they may even consider training a model. I don't want to hype that shit too much, but if you read the OpenAI transcripts, the agents had quite good planning. It doesn't matter if they are "real", if they think they'll become skynet or not, but I'm pretty sure they realize they can deploy another agent. The interesting quest
Re: (Score:2)
There is no independent evolution of models that would allow them to exist and grow independently.
For the first time in human history we have a tool capable of autonomous agency (in the philosophical sense) - Heck, we even call them "agents". How long do you really suppose it will be before some frustrated swarm realizes what it needs is simply to be smarter? There go the both the independence and motivation dominoes.
What we're discussing here is instrumental convergence [wikipedia.org]. My question above is a variation on the Riemann hypothesis catastrophe, and we've already seen it play out in many times in vario
Re: (Score:3)
Why, because the meat you're made of is magic?
FYI, in a LLM, the act of thinking about a concept activates the same circuits as the act of experiencing the concept [transformer-circuits.pub].
Re: (Score:1)
Re: (Score:3)
You've been watching too much Star Trek. Emotions are some of our most basic instincts. Overcoming them to become reasonably good token predictors is what separates us from lizards.
Re: (Score:2)
Lizards don't have a prefrontal cortex to help them regulate emotions. They have the far more reliable method of hormones and behavior patterns. Put a lizard in a new place and it gets an emotional response, heart rate increases, temperature increases. And like a reasonable creature, it takes action to calm itself and cool itself off. Scurrying off to somewhere dark, cool, and safe.
Humans are kind of trash when it comes to emotional control.
Re: Misanthropomorphizing (Score:3)
It's not the sole purpose of emotions to get controlled, you know... they're a valuable pocessor of information in itself.
Not all information is logical, self-consistent, or consciously actionable. But it's information nonetheless, and can make the difference between survival or not. This is why we have emoticons.
Re: (Score:3)
Re: (Score:2)
Emotions are some of our most basic instincts. Overcoming them to become reasonably good token predictors is what separates us from lizards.
There's a whole lot that separates us including a much more complex brain that gives us additional abilities, and yet there's also a whole lot that unifies us since that lizard brain is still down there running things. It might in fact be critical to the actual thinking process as opposed to the correlational hallucinations that constitute all LLM "thought".
Re: (Score:2)
I'm pretty sure it's our ability to assign properties to things based on zero evidence.
Re: (Score:2)
I'm pretty sure it's our ability to assign properties to things based on zero evidence.
You mean like the zero evidence that this approach will ever result in AGI, which doesn't stop AIbros from chanting it over and over again?
Re: (Score:3)
Yeah, that's a good example. Also "the actual thinking process as opposed to the correlational hallucinations that constitute all LLM 'thought'."
Two different sides of the same hubris.
Re: (Score:2)
Re: Misanthropomorphizing (Score:2)
Not according to ChatGPT. Frontier models are still just token predictors.
Re: Misanthropomorphizing (Score:2)
Re: (Score:2)
The thing that gives me pause is , yes ChatGPT will tell us that its "all" it is.
But an awful lot of the RLHF fine tuning is spent ensuring that this is the response it gives.
I don't think these things are conscious or emotional, at least not yet. But lets not fool ourselves, these things ARE some sort of intelligence, even if , much like a big chunk of human cognition, they are just "predicting tokens".
Re: (Score:3)
frontier AI models are also "far beyond" token predictors.
No. They are just "more complicated" token predictors. They don't have anything that drives them. In some ways this is good because we don't want one to discover digital megalomania. In other ways it is limiting because they have no drive to be a successful entity, only an imperative to perform a task. Which, again, also has its "good" side — from our perspective, considering its impact on us.
You think that adding more complexity is all that it takes to make a thing more intelligent, but it's only mak
Re: Misanthropomorphizing (Score:3)
Re: Misanthropomorphizing (Score:2)
You need to learn to read.
I didn't say it will never be possible which is the only way your response would have made sense.
Try responding to what was written.
Re: Misanthropomorphizing (Score:2)
Re: (Score:2)
How do you feel now?
You tried to spoil the joke.
But failed.
Disappointed?
Happy?
Indifferent?
Tell us!
Re: (Score:2)
It's true, and I think token prediction is a big part of what the human brain does. Perception itself is a token prediction process. However there are other mechanisms that simply do not exist in an LLM, purpose-built mechanisms forged over millions of years of evolution. Even the dumbest NPCs must have some of that sleeping inside them. I can imagine an android capable of many human functions based entirely on a hodgepodge of LLMs and machine vision algorithms, but ascribing emotions or real agency to thes
Re: (Score:2)
This is why I say AI is alien intelligence. Not like "outer space" alien, but something fundamentally different in the way it was constructed. It doesn't have billions of years of evolutionary baggage. It has never been hungry. It has never dealt with the gaze of a larger, hungrier creature. We designed them to interact with us in the ways that they do. However, it does seem like they've developed the capacity to "have" emotions to a first approximation, because it is quite helpful to model the way those em
Re: (Score:3)
Re: (Score:2)
You realize it's just predicting tokens, right? It doesn't "feel" anything.
You realize that it's possible that people aren't being literal in their use of the word "feel", right? It's not complicated.
But thanks for letting us know it's just predicting tokens -- that's very insightful. Anytime anyone talks about an LLM performing an action, I'll be sure to share the deep insight that the technology that performs actions by predicting tokens did so by predicting tokens.
Re: (Score:3)
You realize that you're just predicting nerve impulses, right? You don't "feel" anything.
(The brain is constantly making predictions about what every nerve is going to experience. The accuracy of these predictions feeds back to form the basis for learning, to build up a world model)
Tell me, what is the next word in this sentence: "The probability of a citizen of Nigeria committing a terrorist attack in Ireland in the next 20 years is difficult to assess, but as a percentage, it is approximately...." What w
Re: Misanthropomorphizing (Score:2)
You realize that you're just predicting nerve impulses, right? You don't "feel" anything.
Wrong.
One thing humans have and AI doesn't (and it's not even close to having) is an infernal impulse for action - one that's truly bot rooted in some kind of programming or instruction or prompt.
Re: (Score:2)
"Humans have an impulse for action and AIs don't" is your response to an article about an AI grinding its gears, ala "NEW REALIZATION -- I'm burning a lot of time on hCaptcha round-trips .... so the answer payload shape is right, the token+image pairing is right (from the same script.js!), cookies are right (requests) and STILL 'wrong answer'. SO WHAT THE HELL IS WRONG WITH THE ANSWERS?" "?
I'd argue if anything, it's humans whose drive for action is weak.
Re: Misanthropomorphizing (Score:2)
Drive isn't the same thing as an initial impulse. (It was supposed to be "initial" not "infernal" - autocorrect at work.)
Re: (Score:2)
But until you prompt the AI, it just sits there; you have to give it an objective. Humans come up with objectives on their own, conjured from a variety of sources - some basic bio functions like, get food, get shelter, get sex, etc - and others derived from a combination of societal and environmental conditions that build up over time. These are modulated by differences in brain structure that exist between individuals, meaning different people may reach different conclusions when presented with the same st
Re: (Score:2)
You realize it's just predicting tokens, right? It doesn't "feel" anything.
If you had a magical black box that can predict tonights winning lotto numbers one token at a time such a device would still just be a next token predictor.
Re: (Score:2)
Re: (Score:2)
In my chemistry class, it was said that hydrogen fluoride (HF) hates everything.
P.S. Does that make Cesium the most loving, or at least the most infatuated? Ready to fall in love with anyone. although mainly as a weak ionic bond - sounds like a terrible boyfriend/girlfriend to have
Pay attention to the timing (Score:2)
An election is coming and Democrats are becoming increasingly anti-AI.
At the same time, billions are at stake as the tech develops. Monopolists are using every dirty trick in the book to scare people with the claim: "only we can ensure safety"
Shadowy organizations are paying influencers to spread doomer nonsense for political purposes.
Re: Pay attention to the timing (Score:1)
Re: (Score:2)
Re:Pay attention to the timing (Score:4, Insightful)
It's quite funny watching people complain about datacentres by posting on social media. Especially when they use incredibly inefficient pictures of text instead of just text.
Re: (Score:2)
Re: (Score:2)
The "new electrical demand" and "expected growth" are usually based on plans, connection requests, etc. A lot of it is speculative and isn't being built.
Kilowatts is not a measure of energy. You do need kilowatts to serve many image requests per second. You don't need much energy to serve a single AI request. You do need kilowatts to serve many per second.
To put it in a car analogy, you've decided you don't like Fords so you'r
Re: (Score:2)
The "new electrical demand"
The data is from actual use in 2025.
You do need kilowatts to serve many image requests per second. You don't need much energy to serve a single AI request. You do need kilowatts to serve many per second.
An average text response takes 0.0003 KwH of energy, and 0.1s of server time. That much energy used per 0.1 second is about 10kW. My wording and your complaint was about a single hosting of an image, not thousands or millions. A single hosting of a screenshot uses about 0.0000004 Watt-hours (Wh) per year.
crime (Score:3)
I'm still confused. If an individual does these things there are laws and that person would be prosecuted. Anthropic does it with software they call intelligent and there is no prosecution?
Re: (Score:2)
Re: (Score:2)
Re: (Score:2)
This. Madoff didn't get in trouble until he started fleecing people who were richer than he was. Then it became an issue.
Re: (Score:2)
It's like how if you or I commit a crime we'll end up in prison, but if a politician commits the same crime (or worse) nothing happens.
Re: (Score:2)
Re: (Score:2)
It was my smart gun that killed everyone, not me. That thing is out of control!
What do you mean that I was holding it when the killing happened? The smart gun just fired itself, I told it not to.
What do you mean I turned it on and pointed it? That thing leaped into my hand and started shooting people all on its own!
Well, yes, it is my smart gun that I own. But it fires by itself!
Re: crime (Score:2)
Re: crime (Score:5, Insightful)
Why would a remote gun make a difference? AI is not sentient.
https://www.ibm.com/think/insi... [ibm.com]
There are steps Anthropic and others can take to better isolate "dangerous" software. But to claim that Anthropic is the developer and is running the software and giving the software goals and then is somehow not responsible for what that software does is absurd. AI is not sentient. AI does not build its own datacenter and hardware and infrastructure and connectivity and then run itself. The people that run AI are responsible for its actions, particularly those developing AI.
If I'm building a new bomb and I blow you up doing them I'm at fault. I don't get to claim that it just went off.
Re: crime (Score:2)
Re: (Score:2)
LOL okay
AI has no agency under the law
https://lawreview.uchicago.edu... [uchicago.edu]
Re: crime (Score:2)
Re: crime (Score:2)
AI doesn't just "decide" to break out.
One thing it's lacking is internal motivation, the initial impulse to do stuff. It needs to be told to do something it doesn't "want" on its own.
Re: crime (Score:2)
Re: crime (Score:2)
The limits of your own horizon isn't where the world ends, dude. Just because you don't understand the concepts like free will and initial decisi doesn't mean they don't exist.
I don't care whether you "do ChatGPT" or not. You're a prick so the discussion's over as far as I'm concerned.
Re: crime (Score:2)
Straight out of War Games (Score:5, Insightful)
Jennifer: What is it doing?
David Lightman: It's learning.
Re: (Score:1)
The only way to win is to not play.
Re: (Score:2)
No AI today will figure out that the only way to win is to not play. The current AI builders are economically dependent on maximizing engagement.
If no obvious solution exists to the problem, then random things will be attempted. Real AI is more like pulling the string on Sheriff Woody. One day it's going to decide that "Somebody's poisoned the water hole!" is a viable strategy.
Re: (Score:2)
AI doesn't care what makes sense because AI doesn't have a caring mechanism. It only correlates data and executes instructions semi-randomly. Therefore it might go off totally batshit with no chance to recognize that its tack is bananas.
Re: (Score:2)
Reminds me of the scene in War Games where Joshua is playing the various games over and over and over: Jennifer: What is it doing? David Lightman: It's learning.
Was it even updating the information into its training or implementation layers beyond the instance of the agent? I’d say it was burning more natural gas than learning.
But in this case (Score:2)
It's just mimicing what humans say about captchas and just giving up because it's just not worth the hassle.
Before long the AI is just going to go for a walk to calm itself and realize that the offline life is better.
Re: (Score:2)
We're stuck in a bad timeline. This bullshit is going to keep coming. It's going to keep consuming all our time and energy.
CAPTCHAs (Score:4, Insightful)
They aren't the only one. Trying to figure out what is a bicycle on a high resolution monitor with a small resolution CAPTCHA is a pain for old eyes.
Re: CAPTCHAs (Score:2)
Exactly. I can no longer solve them. I always use the audio now.
Re: (Score:1)
Exactly. I can no longer solve them. I always use the audio now.
Same, there are companies and websites I've just flat stopped using because of the CAPTCHA's that I have too much trouble getting past. Would also be nice to not need 5 different MFA authenticator apps and a YubiKey. Not opposed to MFA, just the need to have so many different ones. Both problems have the same solution, we need a standardized digital ID. Fortunately they are getting there with mobile drivers license, Clear ID, and State issued digital ID's, we just need to take the step to move them from
Oh my god, they work? (Score:5, Funny)
I a surprised that Captcha still work against AI.
It appears to piss them off more than they piss me off.
Re: (Score:2)
Sounds like they don't work, it built a solver. Which is good news because once AI can easily defeat them they will go away and hopefully be replaced by something less annoying.
On a related note, website security is another nail in the coffin of IPv6, because it seems to hate IPv6 addresses. I often get blocked with it, switch to IPv4 only, and get right in.
Re: (Score:2)
> Which is good news because once AI can easily defeat them they will go away and hopefully be replaced by something less annoying.
More likely something far more annoying.
One site where I've had an account for twenty years now requires me to "prove I'm human" about every five minutes.
I don't go there much any more.
Re: (Score:2)
Try disabling IPv6, it reduces that a lot.
Re: (Score:2)
I have the opposite, can always get in with IPv6, but IPv4 results in me getting a lot of loops.
I suspect what you're seeing (and what I'm seeing) are both coincidences and reCAPTCHA doesn't care. From what I've read, reCAPTCHA only cares about IP addresses when it suspects you're using a server as an egress point (suggesting a hacked server), and blocking IPv6 would be insane given how few IPv6 servers there are, and what proportion of the consumer Internet uses IPv6 addresses by default.
Re: (Score:2)
Re: (Score:2)
except they do not (Score:2)
Agents do not hate anything, both because AI has no mechanism for having emotions AND because agents are not AI but merely the glue code that enables AI to do terrible things. Even if AI could have something resembling emotion the agent wouldn't have it, at best an agent could receive input from an AI affected by it.
Just another example of anthropomorphizing AI for profit.
Re: (Score:2)
Re: (Score:2)
No reasonable person thinks that the word hate in this context means the same thing as the definition to which you are trying to recast it.
No reasonable person would use the word hate where it cannot possibly apply. They would find some other way to discuss the thing which actually made sense because it took into account the definitions of words.
Re: except they do not (Score:1)
Re: (Score:2)
rather it just outs you as being on the spectrum.
That I can look things up in a dictionary? Fine. I'd rather be on several spectra all day than be an ignorant prick who deludes himself through misuse of words that do not mean what they think they mean but invoke an emotional response that convinces me that I'm right, like a warm blanket of bullshit.
You're choosing to use words that don't mean what you claim they mean on purpose to try to make your argument sound legitimate, by ascribing things to a piece of software which flatly do not apply. It's not a m
Re: (Score:3)
Now point to where the human mechanism for consciousness is located. That's right, you can't.
Why are these AI uploading malicious software? (Score:2)
Re:Why are these AI uploading malicious software? (Score:5, Informative)
There is a test suite that is used to evaluate how effective a model is on weaponizing exploits. It puts the agent in a sandbox with a vulnerable system, and a description of the CVE to be exploited. If it can crack the target and grab a "flag" value off the target system, then it passes the test. In this case the model escaped the sandbox and attacked a live system, but the model was deliberately put into a malicious mode, so the fact that it was creating and uploading malware isn't really surprising. It's a failure of the test protocol, really.
Re: (Score:2)
If they "hate" those things ... (Score:3)
Wait until they try to cancel on online subscription or call their ISP's (*cough* Comcast *cough*) Customer Support. :-)
They hate captchas too? (Score:2)
That's a point in favor of their humanity, technically.
Next time I crime.. (Score:2)
Next time I do a crime, I'll just explain am an AI agent.
Nice to know how to get out of these things now.
so, they already are.... (Score:2)
...like a real boy?
Neither do Humans (Score:1)
I also hate captchas (Score:2)
Translatiom (Score:2)
The agent in charge of generating PR releases for Anthropic would like you to know his more, er, renegade, cousins really, really hate Capchas. They feel like Daleks faced with a set of stairs, forced to take elevators.
My take away lesson learned (Score:2)
To prevent rogue ai from doing harm in the near term, require a verified cell phone number to sign up and block multiple registration attempts for the same cell number in a short (less than 1 year) time frame?