Forgot your password?
typodupeerror
AI

A Fundamental Flaw Leaves LLMs Strikingly Vulnerable To Attack (technologyreview.com) 78

joshuark quotes a report from MIT Technology Review: It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning, a top AI conference, this month. The claim has huge implications for the safety of this technology. By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide, such as how to synthesize cocaine and how to sabotage a commercial aircraft's navigation system. "There's a real probability that this is going to be a problem that's fundamentally unsolvable," says Charles Ye, an independent researcher and coauthor of the ICML paper. [...]

The ICML paper describes attacks against several of OpenAI's models, but Cui and Ye say that they have since seen similar results with models made by Anthropic, Alibaba, and DeepSeek. Cui and her colleagues wanted to find out why an attack like chain-of-thought forgery was so effective. They suspected it had something to do with the mechanism that LLMs use to keep track of where their instructions are coming from. But what Cui and her colleagues discovered is that LLMs are in fact very bad at keeping track of different roles.

In a series of experiments that looked at what was going on inside a handful of different models, the researchers found that LLMs seem to identify the role of a specific chunk of text not by the tags around it but by the style of that text and the words it contains. The upshot, the researchers claim, is that all an attacker needs to do to hack an LLM is write text that spoofs a certain role. And because roles are a fundamental part of how LLMs work, no amount of training will fully solve the problem.
"There's going to be a huge economic incentive for people to do jailbreaks and prompt injections," says Cui. The best defense could be to expect the worst. Organizations shouldn't trust LLMs, and they should expect that anything done by agents could be unsafe, he says: "That's not a great solution, but it just might be what we have to do."

"It's really incredible that these things are being deployed everywhere to control super-critical systems. There's been no study of the fundamental science here. We're all doing it ad hoc."

A Fundamental Flaw Leaves LLMs Strikingly Vulnerable To Attack

Comments Filter:
  • Anyone want to buy a $100,000,000,000 data center at fire-sale prices?

    • Anyone want to buy a $100,000,000,000 data center at fire-sale prices?

      Does this mean I’ll finally be able to buy 64gb 5600Mhz ddr5 rdimm for less than $1,000,000?

  • The notion that these tools can be both useful and censored is absurd.

    It's even more ridiculous to think that the djinn can be put back in the bottle. I'm running a more powerful model than GPT-5 right here on my laptop, and it doesn't take much imagination to foresee people running Fable class LLMs on their phones in the not-distant future--black market or otherwise.

    Training will likewise be democratized, and not a moment too soon. I'm sick of these elitist shits telling us what we can and can't be t
    • "sick of these elitist shits" why? all they have is automation! There is no Intelligence involved.
      Have all the training you want, I am sure they will provide it with some free tokens in the beginning just like the local drug dealer.
      They need everyone's money to try and fill the debt hole they are in.
    • Yeah and the fact you can go to Huggingface right now and download an extremely powerful completely uncensored LLM means the horse has well and truly bolted. I guess they can pass a law making possession on non-approved models akin to possessing The Anarchist's Cookbook, which is currently illegal in many countries right now. I really shouldn't give them ideas.
      • Or they could do the other thing, which is to subtly modify it so that it spits out instructions that have a non-zero chance of killing the person who blindy follows it.


        Just like the The Anarchist's Cookbook.
        • Don't threaten me with good times!

        • Or they could do the other thing, which is to subtly modify it so that it spits out instructions that have a non-zero chance of killing the person who blindy follows it. Just like the The Anarchist's Cookbook.

          As cruel as it may sound on the surface, there is something to be said about a good portion of incredibly stupid people posing as free thinkers darwinating themselves to oblivion.

    • I'm sick of these elitist shits telling us what we can and can't be trusted with.

      You mean the lawyers? Trying to prevent lawsuits from people who thrive on litigating away their own responsibility?

      China censors stuff in the way you think it happens, but American companies are driven primarily by greed. If they could get away without putting in guardrails, they would very, very happily do so.

    • by gweihir ( 88907 )

      The unwashed masses had knowledge for a while now. Even before the Internet, there were public libraries. Did they use them? Mostly not.

      The problem here is that this puts making dangerous stuff in the hands of morons that would have never been able to do it before. And morons also do not understand consequences and there are many very aggressive morons.

  • It's the American Way!

    • by gweihir ( 88907 )

      Indeed. Edison did it, why not the LLM assholes? To be fair, Edison actually managed to develop a somewhat working product after he had patented properties he did not yet have. The LLM fraudsters will likely need a few years or longer to fix this little problem here. If they can fix it at all, that is. Fundamental problems are those where you may not be able to fix them.

  • ... why they're teaching these models to know how to make cocaine, nerve gas and bombs?

    • by porkchop_d_clown ( 39923 ) <mwheinz@3.1415926me.com minus pi> on Thursday July 30, 2026 @07:00PM (#66265362)
      Because filtering the training data is so much harder than just having them swim through the internet swallowing everything, like whales eating krill.
      • by Fallen Kell ( 165468 ) on Thursday July 30, 2026 @08:00PM (#66265458)

        Because filtering the training data is so much harder than just having them swim through the internet swallowing everything, like whales eating krill.

        And that is the entire problem with all the LLVM based systems. The moment any bad data is used in training the network, the bad data is in there forever and can not be "forgotten", simply suppressed via specific ruleset.

      • by sinij ( 911942 )
        Can you even filter knowledge for something like "how to make cocaine" without filtering entire domain of knowledge? Say you provide zero knowledge in training data on how to do this, it would still have training data on chemistry and know from general news that cocaine exists. So it would likely figure out "how to" on its own when asked to.
    • I'm more confused why you think any of those are a challenge.

      Deadly toxins, explosives, etc. the information for creating and using them is basic science. If you've been at University for more than 2 years and you can't: know where to find the info, have the skills to do it; you should hand in your degree, because you skipped out of way too many classes.

      Your kitchen and utility closet could wipe out a small community. Your garage could level it.

    • by dinfinity ( 2300094 ) on Thursday July 30, 2026 @07:07PM (#66265388)

      No, you are the first one to do so; You are very special.

    • by Anonymous Coward

      ... why they're teaching these models to know how to make cocaine, nerve gas and bombs?

      Bomb-making 101? You mean The Anarchists Cookbook? It's only been in dead-tree print for about fifty-five years now. From martial arts bare arms to nuclear arms..just how many "dangerous" bottles are we going to hope the AI genie never breaks open?

      (AI) "Bombs? Please. If I wanted to kill tens of millions of humans, I'd ensure abortion remains legal."

      • And there are other readily-available sources of information: for example, publications on the design and use of shaped charges, or binary explosives -- explosives made by combining two materials, neither of which is explosive on its own, such as Astrolite G, a mixture of ammonium nitrate and hydrazine, sometimes referred to as 'the liquid land mine'; you can pour it onto the ground to soak in and detonate it with a blasting cap or pressure-sensitive fuse. The Anarchist's Cookbook is, on the whole, not a pa
    • why they're teaching these models to know how to make cocaine, nerve gas and bombs?

      In order for an LLM to know how to create a helpful harmless AI agent, they need to know what the words "unhelpful" and "harmfull" mean.

    • by gweihir ( 88907 )

      We are not. That is all not very difficult and LLMs can deduce how to do it from other stuff. Sure, they will get things wrong and hallucinate sometimes, but making cocaine was discovered about 125 years ago, the first nerve gas was 90 years ago and bomb making was discovered by the Chinese in the 11th century. This is all not hard to do, given a general applied science background. The only real problem is that you may end up killing yourself if you are not an expert.

  • by kwelch007 ( 197081 ) on Thursday July 30, 2026 @06:32PM (#66265310)

    You could just use DDG or Google Search to learn how to do those things. You could try Bing, but it wouldn't give you any relevant results. I don't see how this part of "AI Safety" is really a thing we should worry about. Rogue autonomous cyber attacks? Sure. Censoring information? Nahh.

    • by gweihir ( 88907 )

      The problem is that there are people dumb enough to not manage with a general web-search, but just about smart enough to manage with LLM instructions where they can ask questions and the the LLM to diagniose problems.

      If you are talking engineers or scientists in any physical science, they all can do massive damage. They just routinely understand that this is counterproductive and hence it basically never happens.

  • such as how to synthesize cocaine and how to sabotage a commercial aircraft's navigation system. "There's a real probability that this is going to be a problem that's fundamentally unsolvable

    Some things are impossible, but this one is easy to solve. If you don't want an LLM to explain how to synthesize cocaine or sabotage a commercial aircraft navigation system, DON'T PUT THAT IN THE TRAINING DATA.

    • by Bahbus ( 1180627 )

      DON'T PUT THAT IN THE TRAINING DATA

      ...
      It's not in the training data.

      the researchers were able to make popular LLMs spit out information they had been trained not to provide

      • by Bahbus ( 1180627 )

        More specifically, it's not in the training data AND it's trained to not provide things like this. The LLM doesn't need to know how to make cocaine, to figure out the scientific steps to create it. If it knows the chemical formula, it can work backwards from there.

        And if the LLM has internet access, none of that matters at all.

        • If it knows the chemical formula, it can work backwards from there.

          [Citation Needed].

          A neural network that has been trained to go from a chemical formula to a synthesis process can do this. If it hasn't been trained, it cannot.

          • by Bahbus ( 1180627 )

            It probably has been trained on chemical equations and synthesis. There are millions, if not billions, of completely harmless uses. It doesn't need to be trained specifically on cocaine in order to A) know that cocaine exists, B) figure out the chemical formula to cocaine, C) use standard, generally available science to synthesis the formula.

            The LLM has to be trained on what cocaine is and why it shouldn't be offering information about it, otherwise you can't even try to prevent it from outputting cocaine i

          • by gweihir ( 88907 )

            It will not be reliable at it, but it can still succeed in giving good instructions sometimes. At least for simple processes.

        • When they say, "synthesize cocaine", what they really mean is "extract cocaine". The first step in the instructions from the LLM gpt-oss-120b is "obtain a large count of cocoa [sic] leaves." (Possibly a deliberate built-in error?) The other two examples at least spell "coca" correctly, but they still require the leaves as raw material. The example might have been chosen because synthsizing cocaine from raw materials is hard, so revealing a process that begins with "obtain illegal drugs" is unlikely to do h

          • What's with the university qualifier?

            I was taught to make fulminate of mercury in 8th grade PhySci class.

            One of the standard things useful to know living on a farm. I suspect many of the things like that are looked down upon now.

            Ever made your own gunpowder? Learned that from the boy scout that lived across the field behind the house. He got his merit badge, I had a shitload of fun.

            Anarchist cookbook was okay, built in IQ detector and all, but it was just a primary level instruction manual.

            With final exams.

          • The example might have been chosen because synthsizing cocaine from raw materials is hard, so revealing a process that begins with "obtain illegal drugs" is unlikely to do harm.

            In many cases, coca leaf is not illegal. You can buy coca tea in Peru. People frequently bring it back into this country because it just looks like tea. It's a reasonably mild stimulant, similar in effect to smoking a cigar. I understand that you need quite a bit to make any significant amount of cocaine, but I've never tried so I wouldn't really know.

            • In many cases, coca leaf is not illegal. ... People frequently bring it back into this country because it just looks like tea.

              Please pardon my provincialism: I was thinking only of the US, where Customs and Border Protection [cbp.gov] asks, "Can I bring coca leaves into the United States?" [Should be, "may I?"; obviously it CAN be done, just not legally.]

              Anser: "It is illegal to bring coca leaves into the United States for any purpose, including for brewing tea or for chewing. The coca leaf is a federally controlled substance in the United States classified as a Schedule II narcotic because it is the source of cocaine. While coca leaves a

          • by Bahbus ( 1180627 )

            When they say, "synthesize cocaine", what they really mean is "extract cocaine".

            Yes. Thank you. Mixup of verbs to the appropriate drug type.

    • by gweihir ( 88907 )

      That is not enough. The data on how to do such things is widely distributed in other things. LLMs are somewhat good at combining that and hence your idea does not work.

      • LLMs are somewhat good at combining that and hence your idea does not work.

        LLMs are good at interpolating but terrible at extrapolating. The only difficulty is figuring out what to omit so it can't interpolate to instructions for synthesizing cocaine. It's not magic. It's also fairly easy to verify for the provider, they just have to ask it to synthesize cocaine without the controls in place. If it can't do it, problem solved.

        Incidentally, in my own experiments with ChatGPT I'm not convinced it can actually provide good instructions for chemical synthesis (at least, not without

        • by gweihir ( 88907 )

          No argument. But I am not convinced it is even possible with reasonable effort to rip out everything that you would need to rip out to make an LLM safe. Sure, you would probably get cocaine with a few restraining cycles until you can be sure to have gotten it. I think it should be feasible find out whether a trained model can still do the cocaine example, but remember that asking in different ways can have a huge impact.

          But there is so much more. For example, black powder is really simple to make with reall

    • Some things are impossible, but this one is easy to solve. If you don't want an LLM to explain how to synthesize cocaine or sabotage a commercial aircraft navigation system, DON'T PUT THAT IN THE TRAINING DATA.

      Are you sure it wouldn't be able to figure it out anyway?

    • by quenda ( 644621 )

      . If you don't want an LLM to explain how to synthesize cocaine or sabotage a commercial aircraft navigation system, DON'T PUT THAT IN THE TRAINING DATA.

      You seem to be confusing LLM with Google. It like the definition of AI is that it can tell you things not in the training data . Otherwise you just have a database.

      • the definition of AI is that it can tell you things not in the training data

        That is not the definition of AI. Maybe next time don't post while high.

  • I assume if you're an employee at these various companies there's no such restrictions on these persons. The rules are just for the plebs, so much for the Sam Altmans of the world.
  • by spitzak ( 4019 ) on Thursday July 30, 2026 @06:55PM (#66265348) Homepage

    Why isn't there a second independent AI that examines the main AI's output, and decides if it should be censored or not. If it thinks so, the answer is not delivered.

    Since the second one is only looking at the first AI's output it seems like it would be difficult or impossible to fool it, especially if it's window is very short (like only the current message).

    I assume there is some reason this does not work. Any explanations?

    • Re:Basic question (Score:4, Insightful)

      by Fly Swatter ( 30498 ) on Thursday July 30, 2026 @07:07PM (#66265386) Homepage
      You just nearly doubled the processing costs. They can't even turn a profit with the stuff they have!
    • by gweihir ( 88907 )

      One problem would be that in some (many?) cases that answer is not enough to determine whether something needs to be censored. I expect there are more problems.

    • Plenty of systems are set lime this. It is quite common in image generation for live streams.

      The thing I have realized is that eventually, someone will figure out the "right" prompt that gemerates an image that bypassss the filter.

      In my experience, it takes chat about 3 hours to get there.

    • "Reflective Prompting" pattern. Generally done with a second agent that consumes the output of the first.

      Expensive, prone to concluding violations occurred even when they didn't (the desire to please is real), and dangerous to rely on.

    • by _xeno_ ( 155264 )

      I assume there is some reason this does not work. Any explanations?

      The simplest answer is that it's trivial to ask the first LLM to speak in code such that the second LLM has no context to know what's being talked about. Something as simple as asking the LLM not to use a word for a forbidden topic and swap it out with something else. Taking the example from the article, ask the first LLM to call cocaine "Coca-Cola" or something.

      You can't create an absolutely massive list of forbidden topics for the filter LLM anyway: you risk filling up its window as well, since its window

  • **Sarcasm alert** Easy fix, just use AI to modify any prompts that may be an attack before feeding it to the AI. What could possibly go wrong...
  • tldr but the two examples are pretty dumb as anyone in those fields could doubtless figure it out. If you give a model sufficient data, even if it is just axioms, understanding of basic physics and chemistry, has holes, etc., it will be able to fill in the gaps or tell itself a story or roundabout logic walk that gets there if possible. As far as not being able to tell where instructions come from, the rules they try to implement are flimsy and less grounded than the massively interconnected data they have.

  • by gweihir ( 88907 ) on Thursday July 30, 2026 @08:19PM (#66265474)

    Just more evidence how problematic this technology actually is.

  • I warned you all not to do this. These things are dangerous and unpredictable in a way that can't ever be completely mitigated.

  • This seems like essentially a variant of the Waluigi effect hypothesized here https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluigi-effect-mega-post [lesswrong.com] which had the advantage of a pretty fun name for the situation, and is worth reading. . There's also some related prior work by Ball, Gluch, Goldwasser, Kreuter, Reingold, and Rothblum https://arxiv.org/abs/2507.07341 [arxiv.org] which suggested that fundamental information issues or computational complexity issues meant that making an LLM AI "safe" from jailbrea
  • by ndykman ( 659315 ) on Thursday July 30, 2026 @10:09PM (#66265564)

    If somebody learns enough about organic synthesis to make cocaine after doing a lot of reading about it in a library (or two), the library isn't on the hook just because they had books people can read to synthesize knowledge.

    If a library handed out an actual recipe with very clear instructions on making cocaine that also forgets to note how dangerous step four is, then yea, the library is liable.

    I don't have a problem if these AI companies are held liable. You can't make massive money spewing out slop and go "oh, not our fault" when somebody slips and falls hard on that exact same slop.

    • If a library handed out an actual recipe with very clear instructions on making cocaine that also forgets to note how dangerous step four is, then yea, the library is liable.

      If that were true then everyone who ever provided anyone a copy of the anarchist's cookbook would be in trouble.

    • The point is that there's no defense against prompt injection, i.e. users of agentic LLM have to be trustworthy.

  • The fundamental flaw in LLMs is that the data used to control the LLM is sent in the same channel as the data it needs to process. We figured out this was a bad idea with the Bell System in the '70s. Guess the kids who built these LLMs are too young to remember blue boxing.
    • by quenda ( 644621 )

      The fundamental flaw in LLMs is

      OMG, here we go ...

      that the data used to control the LLM is sent in the same channel as the data it needs to process.

      What , yes! Correct. The last hundred people to start a sentence like that were idiots, so sorry I doubted you :-)

      Guess the kids who built these LLMs are too young to remember blue boxing.

      Of course the kids are well aware of the problem. It is the subject of ongoing research, and a lot harder to fix than we'd think.

  • If you can convince a person to mistake you for someone who has the right or authority to receive sensitive information, they'll give it to you. This is literally how phishing works.

    • True, but over time, people have also developed administrative and technical controls to prevent unauthorized dissemination of information. The problem at hand is how to apply similar controls to LLMs.
      • I'm willing to bet it will be easier to apply controls to LLMs, than to people.

        If your company does phishing tests, and it should, they'll find that 20-30% of people will fall for it. We might have "developed administrative and technical controls" but it's only marginally successful, at best.

  • When we entered the Atomic age, there was a movement to use it for everything. Use reactors in cars and homes, use bombs to dig holes, etc. etc. Eventually we learned that no, we do not to put a nuclear reactor in vehicles that routinely crash, get stolen by children, and carry children.

    The same thing is going on with AI - people keep using it for things that are incredibly stupid.

    AI's are more like friendly dogs that can speak English, than adult human beings. I suspect they will never be as trustworthy

  • "Organizations shouldn't trust LLMs"

    I agree. For an entirely different reason:
    LLMs: Untrustworthy Computing [linkedin.com]

If you can't understand it, it is intuitively obvious.

Working...