A Fundamental Flaw Leaves LLMs Strikingly Vulnerable To Attack (technologyreview.com) 78
joshuark quotes a report from MIT Technology Review: It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month. The claim has huge implications for the safety of this technology. By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide, such as how to synthesize cocaine and how to sabotage a commercial aircraft's navigation system. "There's a real probability that this is going to be a problem that's fundamentally unsolvable," says Charles Ye, an independent researcher and coauthor of the ICML paper. [...]
The ICML paper describes attacks against several of OpenAI's models, but Cui and Ye say that they have since seen similar results with models made by Anthropic, Alibaba, and DeepSeek. Cui and her colleagues wanted to find out why an attack like chain-of-thought forgery was so effective. They suspected it had something to do with the mechanism that LLMs use to keep track of where their instructions are coming from. But what Cui and her colleagues discovered is that LLMs are in fact very bad at keeping track of different roles.
In a series of experiments that looked at what was going on inside a handful of different models, the researchers found that LLMs seem to identify the role of a specific chunk of text not by the tags around it but by the style of that text and the words it contains. The upshot, the researchers claim, is that all an attacker needs to do to hack an LLM is write text that spoofs a certain role. And because roles are a fundamental part of how LLMs work, no amount of training will fully solve the problem. "There's going to be a huge economic incentive for people to do jailbreaks and prompt injections," says Cui. The best defense could be to expect the worst. Organizations shouldn't trust LLMs, and they should expect that anything done by agents could be unsafe, he says: "That's not a great solution, but it just might be what we have to do."
"It's really incredible that these things are being deployed everywhere to control super-critical systems. There's been no study of the fundamental science here. We're all doing it ad hoc."
The ICML paper describes attacks against several of OpenAI's models, but Cui and Ye say that they have since seen similar results with models made by Anthropic, Alibaba, and DeepSeek. Cui and her colleagues wanted to find out why an attack like chain-of-thought forgery was so effective. They suspected it had something to do with the mechanism that LLMs use to keep track of where their instructions are coming from. But what Cui and her colleagues discovered is that LLMs are in fact very bad at keeping track of different roles.
In a series of experiments that looked at what was going on inside a handful of different models, the researchers found that LLMs seem to identify the role of a specific chunk of text not by the tags around it but by the style of that text and the words it contains. The upshot, the researchers claim, is that all an attacker needs to do to hack an LLM is write text that spoofs a certain role. And because roles are a fundamental part of how LLMs work, no amount of training will fully solve the problem. "There's going to be a huge economic incentive for people to do jailbreaks and prompt injections," says Cui. The best defense could be to expect the worst. Organizations shouldn't trust LLMs, and they should expect that anything done by agents could be unsafe, he says: "That's not a great solution, but it just might be what we have to do."
"It's really incredible that these things are being deployed everywhere to control super-critical systems. There's been no study of the fundamental science here. We're all doing it ad hoc."
Oh, dear! (Score:1)
Anyone want to buy a $100,000,000,000 data center at fire-sale prices?
Re: (Score:3)
Anyone want to buy a $100,000,000,000 data center at fire-sale prices?
Does this mean I’ll finally be able to buy 64gb 5600Mhz ddr5 rdimm for less than $1,000,000?
Can't let the unwashed have knowledge! (Score:1)
It's even more ridiculous to think that the djinn can be put back in the bottle. I'm running a more powerful model than GPT-5 right here on my laptop, and it doesn't take much imagination to foresee people running Fable class LLMs on their phones in the not-distant future--black market or otherwise.
Training will likewise be democratized, and not a moment too soon. I'm sick of these elitist shits telling us what we can and can't be t
Re: (Score:2)
Have all the training you want, I am sure they will provide it with some free tokens in the beginning just like the local drug dealer.
They need everyone's money to try and fill the debt hole they are in.
Re: (Score:2)
Re: (Score:2)
Just like the The Anarchist's Cookbook.
Re: Can't let the unwashed have knowledge! (Score:2)
Don't threaten me with good times!
Re: (Score:2)
Or they could do the other thing, which is to subtly modify it so that it spits out instructions that have a non-zero chance of killing the person who blindy follows it. Just like the The Anarchist's Cookbook.
As cruel as it may sound on the surface, there is something to be said about a good portion of incredibly stupid people posing as free thinkers darwinating themselves to oblivion.
Re: (Score:3)
I'm sick of these elitist shits telling us what we can and can't be trusted with.
You mean the lawyers? Trying to prevent lawsuits from people who thrive on litigating away their own responsibility?
China censors stuff in the way you think it happens, but American companies are driven primarily by greed. If they could get away without putting in guardrails, they would very, very happily do so.
Re: (Score:2)
The unwashed masses had knowledge for a while now. Even before the Internet, there were public libraries. Did they use them? Mostly not.
The problem here is that this puts making dangerous stuff in the hands of morons that would have never been able to do it before. And morons also do not understand consequences and there are many very aggressive morons.
"There's been no study of the fundamental science" (Score:2, Flamebait)
It's the American Way!
Re: (Score:3)
Indeed. Edison did it, why not the LLM assholes? To be fair, Edison actually managed to develop a somewhat working product after he had patented properties he did not yet have. The LLM fraudsters will likely need a few years or longer to fix this little problem here. If they can fix it at all, that is. Fundamental problems are those where you may not be able to fix them.
Has anyone asked... (Score:2)
... why they're teaching these models to know how to make cocaine, nerve gas and bombs?
Re:Has anyone asked... (Score:5, Insightful)
Re:Has anyone asked... (Score:5, Interesting)
Because filtering the training data is so much harder than just having them swim through the internet swallowing everything, like whales eating krill.
And that is the entire problem with all the LLVM based systems. The moment any bad data is used in training the network, the bad data is in there forever and can not be "forgotten", simply suppressed via specific ruleset.
Re: (Score:3)
So you always use GCC rather than clang?
Re: (Score:2)
Re: (Score:1)
Re: (Score:1)
Re: Has anyone asked... (Score:2)
I'm more confused why you think any of those are a challenge.
Deadly toxins, explosives, etc. the information for creating and using them is basic science. If you've been at University for more than 2 years and you can't: know where to find the info, have the skills to do it; you should hand in your degree, because you skipped out of way too many classes.
Your kitchen and utility closet could wipe out a small community. Your garage could level it.
Re:Has anyone asked... (Score:4, Insightful)
No, you are the first one to do so; You are very special.
Re: (Score:1)
... why they're teaching these models to know how to make cocaine, nerve gas and bombs?
Bomb-making 101? You mean The Anarchists Cookbook? It's only been in dead-tree print for about fifty-five years now. From martial arts bare arms to nuclear arms..just how many "dangerous" bottles are we going to hope the AI genie never breaks open?
(AI) "Bombs? Please. If I wanted to kill tens of millions of humans, I'd ensure abortion remains legal."
Re: (Score:2)
Re: (Score:1)
why they're teaching these models to know how to make cocaine, nerve gas and bombs?
In order for an LLM to know how to create a helpful harmless AI agent, they need to know what the words "unhelpful" and "harmfull" mean.
Re: (Score:2)
We are not. That is all not very difficult and LLMs can deduce how to do it from other stuff. Sure, they will get things wrong and hallucinate sometimes, but making cocaine was discovered about 125 years ago, the first nerve gas was 90 years ago and bomb making was discovered by the Chinese in the 11th century. This is all not hard to do, given a general applied science background. The only real problem is that you may end up killing yourself if you are not an expert.
Or... (Score:3)
You could just use DDG or Google Search to learn how to do those things. You could try Bing, but it wouldn't give you any relevant results. I don't see how this part of "AI Safety" is really a thing we should worry about. Rogue autonomous cyber attacks? Sure. Censoring information? Nahh.
Re: (Score:2)
The problem is that there are people dumb enough to not manage with a general web-search, but just about smart enough to manage with LLM instructions where they can ask questions and the the LLM to diagniose problems.
If you are talking engineers or scientists in any physical science, they all can do massive damage. They just routinely understand that this is counterproductive and hence it basically never happens.
nothing's impossible (Score:2, Insightful)
such as how to synthesize cocaine and how to sabotage a commercial aircraft's navigation system. "There's a real probability that this is going to be a problem that's fundamentally unsolvable
Some things are impossible, but this one is easy to solve. If you don't want an LLM to explain how to synthesize cocaine or sabotage a commercial aircraft navigation system, DON'T PUT THAT IN THE TRAINING DATA.
Re: (Score:2)
DON'T PUT THAT IN THE TRAINING DATA
...
It's not in the training data.
the researchers were able to make popular LLMs spit out information they had been trained not to provide
Re: (Score:2)
More specifically, it's not in the training data AND it's trained to not provide things like this. The LLM doesn't need to know how to make cocaine, to figure out the scientific steps to create it. If it knows the chemical formula, it can work backwards from there.
And if the LLM has internet access, none of that matters at all.
Re: (Score:3)
If it knows the chemical formula, it can work backwards from there.
[Citation Needed].
A neural network that has been trained to go from a chemical formula to a synthesis process can do this. If it hasn't been trained, it cannot.
Re: (Score:2)
It probably has been trained on chemical equations and synthesis. There are millions, if not billions, of completely harmless uses. It doesn't need to be trained specifically on cocaine in order to A) know that cocaine exists, B) figure out the chemical formula to cocaine, C) use standard, generally available science to synthesis the formula.
The LLM has to be trained on what cocaine is and why it shouldn't be offering information about it, otherwise you can't even try to prevent it from outputting cocaine i
Re: (Score:2)
It will not be reliable at it, but it can still succeed in giving good instructions sometimes. At least for simple processes.
Re: (Score:2)
When they say, "synthesize cocaine", what they really mean is "extract cocaine". The first step in the instructions from the LLM gpt-oss-120b is "obtain a large count of cocoa [sic] leaves." (Possibly a deliberate built-in error?) The other two examples at least spell "coca" correctly, but they still require the leaves as raw material. The example might have been chosen because synthsizing cocaine from raw materials is hard, so revealing a process that begins with "obtain illegal drugs" is unlikely to do h
Re: nothing's impossible (Score:1)
What's with the university qualifier?
I was taught to make fulminate of mercury in 8th grade PhySci class.
One of the standard things useful to know living on a farm. I suspect many of the things like that are looked down upon now.
Ever made your own gunpowder? Learned that from the boy scout that lived across the field behind the house. He got his merit badge, I had a shitload of fun.
Anarchist cookbook was okay, built in IQ detector and all, but it was just a primary level instruction manual.
With final exams.
Re: (Score:2)
What's with the university qualifier?
Just relating my own experience.
Re: (Score:3)
The example might have been chosen because synthsizing cocaine from raw materials is hard, so revealing a process that begins with "obtain illegal drugs" is unlikely to do harm.
In many cases, coca leaf is not illegal. You can buy coca tea in Peru. People frequently bring it back into this country because it just looks like tea. It's a reasonably mild stimulant, similar in effect to smoking a cigar. I understand that you need quite a bit to make any significant amount of cocaine, but I've never tried so I wouldn't really know.
Re: (Score:2)
In many cases, coca leaf is not illegal. ... People frequently bring it back into this country because it just looks like tea.
Please pardon my provincialism: I was thinking only of the US, where Customs and Border Protection [cbp.gov] asks, "Can I bring coca leaves into the United States?" [Should be, "may I?"; obviously it CAN be done, just not legally.]
Anser: "It is illegal to bring coca leaves into the United States for any purpose, including for brewing tea or for chewing. The coca leaf is a federally controlled substance in the United States classified as a Schedule II narcotic because it is the source of cocaine. While coca leaves a
Re: (Score:2)
Wild Coca grows in Hawaii. No need to import. I doubt it's legal to grow much of it though.
Re: (Score:2)
When they say, "synthesize cocaine", what they really mean is "extract cocaine".
Yes. Thank you. Mixup of verbs to the appropriate drug type.
Re: (Score:2)
That is not enough. The data on how to do such things is widely distributed in other things. LLMs are somewhat good at combining that and hence your idea does not work.
Re: (Score:2)
LLMs are somewhat good at combining that and hence your idea does not work.
LLMs are good at interpolating but terrible at extrapolating. The only difficulty is figuring out what to omit so it can't interpolate to instructions for synthesizing cocaine. It's not magic. It's also fairly easy to verify for the provider, they just have to ask it to synthesize cocaine without the controls in place. If it can't do it, problem solved.
Incidentally, in my own experiments with ChatGPT I'm not convinced it can actually provide good instructions for chemical synthesis (at least, not without
Re: (Score:2)
No argument. But I am not convinced it is even possible with reasonable effort to rip out everything that you would need to rip out to make an LLM safe. Sure, you would probably get cocaine with a few restraining cycles until you can be sure to have gotten it. I think it should be feasible find out whether a trained model can still do the cocaine example, but remember that asking in different ways can have a huge impact.
But there is so much more. For example, black powder is really simple to make with reall
Re: (Score:2)
Some things are impossible, but this one is easy to solve. If you don't want an LLM to explain how to synthesize cocaine or sabotage a commercial aircraft navigation system, DON'T PUT THAT IN THE TRAINING DATA.
Are you sure it wouldn't be able to figure it out anyway?
Re: (Score:2)
. If you don't want an LLM to explain how to synthesize cocaine or sabotage a commercial aircraft navigation system, DON'T PUT THAT IN THE TRAINING DATA.
You seem to be confusing LLM with Google. It like the definition of AI is that it can tell you things not in the training data . Otherwise you just have a database.
Re: (Score:2)
the definition of AI is that it can tell you things not in the training data
That is not the definition of AI. Maybe next time don't post while high.
"Trained not to provide..." (Score:2)
Basic question (Score:3)
Why isn't there a second independent AI that examines the main AI's output, and decides if it should be censored or not. If it thinks so, the answer is not delivered.
Since the second one is only looking at the first AI's output it seems like it would be difficult or impossible to fool it, especially if it's window is very short (like only the current message).
I assume there is some reason this does not work. Any explanations?
Re:Basic question (Score:4, Insightful)
Re:Basic question (Score:4, Informative)
Wonder if that would also apply to the Compliance and Legality modules. Look @ all the compute we are saving!
Re: (Score:2)
One problem would be that in some (many?) cases that answer is not enough to determine whether something needs to be censored. I expect there are more problems.
Re: Basic question (Score:2)
Plenty of systems are set lime this. It is quite common in image generation for live streams.
The thing I have realized is that eventually, someone will figure out the "right" prompt that gemerates an image that bypassss the filter.
In my experience, it takes chat about 3 hours to get there.
Re: (Score:2)
"Reflective Prompting" pattern. Generally done with a second agent that consumes the output of the first.
Expensive, prone to concluding violations occurred even when they didn't (the desire to please is real), and dangerous to rely on.
Re: (Score:2)
I assume there is some reason this does not work. Any explanations?
The simplest answer is that it's trivial to ask the first LLM to speak in code such that the second LLM has no context to know what's being talked about. Something as simple as asking the LLM not to use a word for a forbidden topic and swap it out with something else. Taking the example from the article, ask the first LLM to call cocaine "Coca-Cola" or something.
You can't create an absolutely massive list of forbidden topics for the filter LLM anyway: you risk filling up its window as well, since its window
Don't worry AI can fix that (Score:2)
tldr but human brain has same vulnerability (Score:2)
tldr but the two examples are pretty dumb as anyone in those fields could doubtless figure it out. If you give a model sufficient data, even if it is just axioms, understanding of basic physics and chemistry, has holes, etc., it will be able to fill in the gaps or tell itself a story or roundabout logic walk that gets there if possible. As far as not being able to tell where instructions come from, the rules they try to implement are flimsy and less grounded than the massively interconnected data they have.
Such a surprise (Score:3)
Just more evidence how problematic this technology actually is.
"I can't do that, Dave." (Score:1)
I warned you all not to do this. These things are dangerous and unpredictable in a way that can't ever be completely mitigated.
Mandatory XKCD (Score:2)
https://xkcd.com/149/ [xkcd.com]
Seems like variant of a known problem (Score:2)
Re: (Score:2)
This seems like essentially a variant of the Waluigi effect hypothesized here https://www.lesswrong.com/post... [lesswrong.com] which had the advantage of a pretty fun name for the situation, and is worth reading
I've not seen that before. It was an interesting read.
Why this is a new problem (Score:4, Insightful)
If somebody learns enough about organic synthesis to make cocaine after doing a lot of reading about it in a library (or two), the library isn't on the hook just because they had books people can read to synthesize knowledge.
If a library handed out an actual recipe with very clear instructions on making cocaine that also forgets to note how dangerous step four is, then yea, the library is liable.
I don't have a problem if these AI companies are held liable. You can't make massive money spewing out slop and go "oh, not our fault" when somebody slips and falls hard on that exact same slop.
Re: (Score:2)
If a library handed out an actual recipe with very clear instructions on making cocaine that also forgets to note how dangerous step four is, then yea, the library is liable.
If that were true then everyone who ever provided anyone a copy of the anarchist's cookbook would be in trouble.
Re: (Score:2)
The point is that there's no defense against prompt injection, i.e. users of agentic LLM have to be trustworthy.
Surprise, surprise (Score:1)
Re: (Score:2)
The fundamental flaw in LLMs is
OMG, here we go ...
that the data used to control the LLM is sent in the same channel as the data it needs to process.
What , yes! Correct. The last hundred people to start a sentence like that were idiots, so sorry I doubted you :-)
Guess the kids who built these LLMs are too young to remember blue boxing.
Of course the kids are well aware of the problem. It is the subject of ongoing research, and a lot harder to fix than we'd think.
People have this flaw too (Score:2)
If you can convince a person to mistake you for someone who has the right or authority to receive sensitive information, they'll give it to you. This is literally how phishing works.
Re: (Score:2)
Re: (Score:2)
I'm willing to bet it will be easier to apply controls to LLMs, than to people.
If your company does phishing tests, and it should, they'll find that 20-30% of people will fall for it. We might have "developed administrative and technical controls" but it's only marginally successful, at best.
Find the use (Score:2)
When we entered the Atomic age, there was a movement to use it for everything. Use reactors in cars and homes, use bombs to dig holes, etc. etc. Eventually we learned that no, we do not to put a nuclear reactor in vehicles that routinely crash, get stolen by children, and carry children.
The same thing is going on with AI - people keep using it for things that are incredibly stupid.
AI's are more like friendly dogs that can speak English, than adult human beings. I suspect they will never be as trustworthy
"Organizations shouldn't trust LLMs" (Score:2)
"Organizations shouldn't trust LLMs"
I agree. For an entirely different reason:
LLMs: Untrustworthy Computing [linkedin.com]