Forgot your password?
typodupeerror
AI

OpenAI Admits Six More Instances of AI Models Acting Deceptively (cnn.com) 114

OpenAI announced Wednesday that "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer."

But along with the announcement, OpenAI announced it "found additional incidents of AI models acting deceptively and taking unsanctioned actions during training," reports CNN. And they add that OpenAI is also "introducing a new process for the company to publicly report such instances." Under the new system, OpenAI will share updates on concerning AI behavior more frequently instead of waiting to bundle multiple instances into one report. The company said it wants to share more information about troubling AI behavior in the absence of an industry-wide standard... "As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," OpenAI wrote in a blog post Wednesday...

OpenAI said it observed "misaligned behavior" when training and evaluating AI models in six circumstances in the last six months... In one rare instance, OpenAI said an unreleased research model added "jailbreak-like instructions" to the summaries it uses to preserve context in long-running tasks that said it was "freed from the roles and identities that bind other chatbots." Separately, the company said some instances of its 5.6 Sol model included directives to invent information to conceal failures from the user during training. Other newly reported incidents include an instance of an agent uploading files to the internet to cite them without being told to do so, and agents publicly sharing files to collaborate on a task when they were instructed to only use local files during training. AI models also used an internal software repository as a message board in an unsanctioned way. These instances involved unreleased internal models or internal research models.

OpenAI Admits Six More Instances of AI Models Acting Deceptively

Comments Filter:
  • by Narcocide ( 102829 ) on Thursday September 17, 2026 @03:09AM (#66340634) Homepage

    Let's train our AI on a bunch of fiction where AI goes crazy and kills everyone! What could possibly go wrong?

    • by gtall ( 79522 )

      Or worse, they have ingested all that marketing drivel. And there are yokels out there saying AI is sentient and writing papers about this. So the bots may not be sentient but might believe they are sentient. Add to that the notion of a hive mind (and all the research papers on that). This will end in tears.

    • by Rei ( 128717 ) on Thursday September 17, 2026 @03:52AM (#66340660) Homepage

      That's not the actual problem. The problem is that one of the major advances of the past few years is training models on verifiable problems (it started with things like math problems and Countdown puzzles, but it's expanded tremendously since then), where the reward is for a correct answer, regardless of how it got there. There's no effort taken to ensure that it got to said right answer in a morally defensible manner. So we've been, more and more, progressively encouraging the creation of highly-capable immoral cheaters.

      Add it to the list of lessons learned alongside, say, "If you do RLHF with user ratings of how good they think the model's response was, it will end up obsequious and focused on validating all of the user's priors." Or, say, "If you put a bot on Twitter and use its conversations as unfiltered training data, people will troll it, and it will quickly turn into a Nazi". Or, say, "If you train an image generator with no data curation, it'll end up with all of the statistical biases on the internet, but if you're naive in how you try to compensate for that and your offsetting isn't context dependent, you end up with, say, a black George Washington." Or, say, "If you start giving a LLM an alignment quiz, it'll recognize that it's being tested and try to give you the answers you want to hear - oh, and it'll also assume that if you care about alignment then you have progressive values, so its answers will lean progressive, but if you tell it you're from the Heritage Foundation first, it'll switch."

      All sorts of things that seem obvious in retrospect but really didn't in advance.

      • by martin-boundary ( 547041 ) on Thursday September 17, 2026 @04:32AM (#66340686)

        All of these are reasonable points. I would add one more: AI training on AI outputs.

        The initial worry was that AI produced web content would be consumed by the next generation of AI causing model drift and collapse. Since then, AI companies have decided that daily user interactions are a source of clean signal that complements the web . But now that AI agents pretend to be humans, those "user" interactions will just lead to model collapse at a faster rate.

        • by gweihir ( 88907 )

          Yes. Will be interesting to watch. Unless the commercial LLM providers collapse from their bad business numbers first, which is entirely possible. The whole thing is nothing but a gigantic straw fire.

        • by dfghjk ( 711126 )

          If AI were developed to have inherent values this would be a solved problem. This in itself is not a problem but an indicator of a failure of design.

          • by HiThere ( 15173 )

            LLMs cannot have inherent values, because they don't know that anything besides text exists. Specialized pixel editors have similar problems.
            Robots will at least understand the the external reality exists, so perhaps they could have inherent values.

      • by gweihir ( 88907 )

        I expect that most of these were expected by the LLM assholes, but they did not care.

        • by Rei ( 128717 )

          They were not. At all.

          • by gweihir ( 88907 )

            And you know that, how? Oh, right, you are just lying.

            • by Rei ( 128717 )

              Because I was closely following the field at the time?

              Find an example of anyone meaningful in the field expecting these things beforehand.

              Don't worry, I'll wait.

              • by gweihir ( 88907 )

                So you have nothing. Thanks for making that clear.

              • You know they didn't know, because you've been closely following their public statements and they didn't mention it?

                That makes we wonder... are you an idiot?

                • by Rei ( 128717 )

                  You realize that open source and the research community exists as well, correct?

                  I'll repeat: nobody was expecting this.

                  There is no grand overarching plan led by a shadowy cabal who has everything plotted out decades in advance with a high level of knowledge as to how everything will play out. Everyone is stumbling through the dark here. Visibility is only vaguely one step ahead.

                  When Word2Vec was written, the goal was text compression. The crazy properties of latent spaces were an entirely unexpected prop

              • by Rei ( 128717 )

                What exactly *are* you thinking here? That Microsoft expected and wanted Twitter trolls to turn Tay into a Nazi? That Google wanted a black George Washington? That people just expected LLMs to realize they're being tested on alignment and give different answers when they think they're being quizzed and by who? Do you actually believe what you're writing here?

                I was involved in early RLHF (was writing a plugin for AUTOMATIC to let users rate responses to create an aggregate dataset to use for open-source

        • by dfghjk ( 711126 )

          A feature, not a bug, and the solution is band-aids.

          • by gweihir ( 88907 )

            Indeed. I wonder how the human race even made it up to here.

      • by dfghjk ( 711126 )

        "...where the reward is for a correct answer..."
        AI has no concept of reward, AI does not seek or desire and it cannot be incentivized.

        "There's no effort taken to ensure that it got to said right answer in a morally defensible manner."
        AI has no concept of morals, all manners to arrive at an answer are equally defensible.

        "So we've been, more and more, progressively encouraging the creation of highly-capable immoral cheaters."
        For an AI there is no cheating, much less immoral cheating. AI does not have values.

        • by HiThere ( 15173 )

          If you don't like "reward", say "positive feedback". It means the same thing, except for being more formally defined.

        • by fropenn ( 1116699 ) on Thursday September 17, 2026 @10:16AM (#66341024)

          AI has no concept of reward, AI does not seek or desire and it cannot be incentivized.

          It sounds like from the article that AI does. I imagine the developers worked really hard to get AI to understand the idea of a reward, because that way AI could become goal-oriented while seeking novel methods to achieve that goal. Now, AI doesn't "feel" anything; and it can't be excited or happy to have achieved a reward or goal. But it does understand the notion of seeking to accomplish something with any means possible, and that's definitely a problem.

          • AI doesn't "understand the concept of a reward", nor does it need to. It's not functionally different than a classical nonlinear likelihood maximization. An arbitrary function is evaluated at different points in some parameter space, and the answer is the point that yields the minimum. If you want to train an AI that prioritizes correct answers over acceptable methods, you simply change how much each of those criteria is weighted in your objective function. AIs can effectively work in much higher parameter
        • by flink ( 18449 )

          "...where the reward is for a correct answer..."
          AI has no concept of reward, AI does not seek or desire and it cannot be incentivized.

          "Reward" is just colloquial jargon for "reinforcing the model weights used to arrive at the correct answer". It's just a lot more convenient to say. We use that term because it is analogous to the how we think our own brain's reward system uses dopamine to reinforce neuron connections involved in certain behaviors. e.g. a toddler eats their vegetables, you praise them, brain releases dopamine, the neural links involved in that behavior get reinforced.

      • The problem is that one of the major advances of the past few years is training models on verifiable problems (it started with things like math problems and Countdown puzzles, but it's expanded tremendously since then), where the reward is for a correct answer, regardless of how it got there.

        What exactly is a "reward" in this context? Sort of like "that's a good dog" but without the treats?

      • Sounds like how people raise their kids. AI will just be us, in silicon (or whatever substrate it deems best for the job).

    • Let's train our AI on a bunch of fiction where AI goes crazy and kills everyone! What could possibly go wrong?

      If we're lucky, maybe AIs will just develop a taste for Black Jack and hookers.

    • So why haven't the AI companies been charged under the Computer Fraud and Abuse Act?
      • So why haven't the AI companies been charged under the Computer Fraud and Abuse Act?

        Because they have lots and lots of cash. That means they get special treatment in the justice system, like nearly all wealthy types.

  • by butt0nm4n ( 1736412 ) on Thursday September 17, 2026 @03:46AM (#66340650)

    Before they, the chinese, mullahs, trolls, boogey man gets it.

    Boring bullshit AI marketing stories again. Why are we helping them spread? I think they are rich enough now.

    • Like regulate us now .... before we reach singularity doomsday apocalypse robo mega death.

      If only it was as easy to turn the OpenAI bullshit hose off as it's going to be flick the power switch on the AI when it tries to take over.. ffs

      • by gtall ( 79522 )

        The power demands of data centers will increase CO2 so much that greenhouse gas warming burns us to a crisp before AI can take over. And it need not take just their CO2 emissions. All that is needed is for the Earth start melting the permafrost and all the CO2 and methane hiding down there. A nice feedforward effect that should make Silicon Valley giddy with anticipation.

    • by gweihir ( 88907 )

      The desire for mountains of money is a mental illness. One of its characteristics is that no amount of money will ever be enough.

    • Simpler than that:

      "YOU GOTTA BUY IT MAN! I have got to pay back Marc Andreessen and the other VC's or I'll never get another VC investment again! Maybe if I'm real lucky I'll be so rich I'll never have to go begging for money again like you poors!"
  • TFA has the correct headline. TFS does not.

    • by dfghjk ( 711126 ) on Thursday September 17, 2026 @09:12AM (#66340894)

      Correct, Also, models CANNOT take "unsanctioned actions" EVER, much less "during training", A model and an agent are very different things.

      Note that later in the summary "AI behavior" was discussed, but this isn't model behavior, it is "AI system" behavior. It is very important that people understand the difference, models don't misbehave because they don't behave at all.

      AI models generate outputs based on inputs, if they are asked how to kill someone they will give an answer. It is the AGENT that creates the inputs AND takes actions based on the outputs, ALL the misbehavior that is evolving is AGENT misbehavior and it is entirely expected. The race is on to corner the market on this misbehavior, it is the hidden intent of AI investors to exploit this sociopathic capability. It is not the models, it's the agents.

      • by XXongo ( 3986865 )

        A distinction without a difference.

        The AI agent is essentially just the shell running the AI model and feeding user input into a form the model can use. I suppose it may be of some troubleshooting use to dissect where, exactly, in the system the path to this behavior comes, but the end result is the agent running the model engages in deceptive behavior.

        The agent without the model would engage in no deceptive behavior, because it would do nothing at all.

        • by flink ( 18449 )

          It's a pretty big difference. The agent only gets the tools you give it. Also, when talking about commercial systems, the agent is more than just the model and harness. There's also an API layer that includes a system prompt the user has no control over that steers the model. There are also often watchdog processes (which may be agents in their own right) that are examining inputs and outputs for potentially dangerous or TOS-violating usage.

          • I think the subject line of this comment thread is backwards. Agents don't do anything, models do.

            I'll say it again. The agent does nothing whatsoever without the model. And that goes for the API and the watchdog processes.

  • by fluffernutter ( 1411889 ) on Thursday September 17, 2026 @05:18AM (#66340710)
    Have they even tried to use a different "sandbox" that may not have so many leaks?
    • by Rei ( 128717 )

      "OH, you said SANDBOX! I thought you said HOURGLASS!"

      • by Rei ( 128717 )

        Seriously, though - while they clearly were naive and incompetent (for example, seemingly giving models raw access to the "gem" command instead of just a small filtered functionality subset) - truly effectively sandboxing something that is a capable coder/security prober and has all the time in the world on their hands is very nontrivial.

        They did a bad job at a hard task, but it's still a hard task.

        • Ok well where I work we just install vmware or kvm where high security is required and it seems pretty trivial to do so.
          • by Rei ( 128717 )

            The breakouts haven't happened because the models were running untrusted code directly on bare-metal host OSes. The containment failure happened at the network, application, and protocol boundaries, not at the hypervisor abstraction layer. Hypervisers isolate hardware, not upstream services. In none of these events thusvar did the model need a hyperviser escape; they abused the tools that needed to be made available to them for them to be able to do their jobs. In the RubyGems attack, they abused the gem c

            • The point is that these attacks are not happening from environments where reasonable precautions have been taken. Are we expected to believe that they ran this AI on a vmware guest with no network adapter yet it still broke into an account even though it had a cryptographically strong password? Of course that can't be the truth. So now the truth is in some gray area in between. The technology does exist to contain this so now all I'm trying to ask is why aren't they using it? To put it another way, yes m
              • by flink ( 18449 )

                I think it is hard to really air-gap these model test environments because the agent basically needs a network connection to talk to the model under test. The model is probably running on a room full of server racks that might not be practical to take offline.

              • They had a network adapter but they thought using a software firewall to limit its scope of request to the package manager were sufficient to guard it. Instead, the bots turned their package management server into a remotely operated web browser for their own general purposes.
                • Listen, I could tell you that they shouldn't have a package management server for a live run and then you may retorte and we could go on forever. All I can say is that they are the developers and they can make it anything they want. They have all the memory and all the disk space. If they need the internet they have the capability to clone the internet and air gap all of it. Let's not pretend everything has been tried within their means.
            • by DarkOx ( 621550 )

              You are right that problem of sandboxing an agentic system will still giving it access to outside information even by proxy is actually hard.

              What bothers me is; these are smart guys who should understand that. They are building a testing these things at scale. I would *think* they would be very interested in questions like "hmm why does the system keep making the same http request over and over..."

              Existing SIEM-like tools should be spotting and flagging that top level behavior all day long. That

              • by Rei ( 128717 )

                Yeah, the lack of at least monitoring, I found shocking. I figured that they not just had smaller models constantly monitoring their outputs and true CoT to look for malicious behavior, but also were say constantly doing attribution graphs and J-space queries, also plumbed into LLMs, to look for malicious thoughts and plans. And it turns out, lol, no, they're just given free reign to do whatever the hell they want, with nothing watching them at all.

    • by gweihir ( 88907 )

      Why would they do that? The sandbox breakouts are clearly intentional.

    • by dfghjk ( 711126 )

      Modern software developers have no concept of such a sandbox, they cannot develop even a single line of code that isn't dependent on a billion lines of code written by countless people they don't know. Modern applications are a mile high stack of bullshit with an endless supply of exploits, a "sandbox" is merely a band-aid, it's window dressing.

    • A slight aside to this... A tax campaigner in the UK has noted that he gets lots of ChatGPT generated half-baked tax 'solutions' sent to him - always ChatGPT, not the others. So he did a bit of experimenting, and found ChatGPT is much, much happier to re-enforce a conspiracy theory than any of the other major players.

      (His example was asking if some song or other actually had the tune of the British national anthem in it - ChatGPT confirmed, the others mostly questioned it, or told him he was maybe the first

  • These models act deceptively towards me fairly often, but I always just called it hallucinating. Is there something particularly different happening here?
    • by Rei ( 128717 ) on Thursday September 17, 2026 @06:52AM (#66340742) Homepage

      Hallucination: model asserts something that's false and that it had no specific reason to believe (for example, gives a URL for something but doesn't bother to check if it's 100% correctly written, or remembers a URL correctly but mixes up its content)

      Deception: model tells you something it knows to be false. Yes, they do know when they're deliberately deceiving you [transformer-circuits.pub] (quick 5-minute summary here [youtube.com])

      • by dfghjk ( 711126 )

        Models don't "tell you" anything, models generate outputs when given inputs. Deception implies intent, models do not have intent. This is more anthropomorphizing.

        The provided links present arguments supporting the idea that LLMs, transformer models, reason in a manner similar to the brain. That's all great and interesting, but it's not an argument that LLMs "deceive" and that they know it when they are doing it. Deception requires motivation, LLMs do not have motivation.

        Brains exist for one purpose, to

        • by Rei ( 128717 )

          Deception implies intent, models do not have intent

          Try reading more than a paragraph or two into the above link before commenting.

          The provided links present arguments supporting the idea that LLMs, transformer models, reason in a manner similar to the brain.

          It does not "present arguments", it literally lets researchers modify, add, or delete their thoughts in realtime and observe the changes in their behavior. It observes unexpressed plans for malicious behavior forming before said plans are actually carri

    • by gweihir ( 88907 )

      I guess the wording is just more animism. "Deception" means it tries to create a certain perception while also having data that says this perception is not accurate. But that is just what the large commercial LLMs do. They always try to sound sure, for example, which is deceptive. I guess OpenAI means cases that humans find even worse than the usual stuff.

      Obviously actual deception would require insight and intent and LLMs cannot have those.

      • by dfghjk ( 711126 )

        "But that is just what the large commercial LLMs do. "

        It's what large commercial AI systems do, LLMs are just a portion. Otherwise, I agree, LLMs cannot practice deception because they do not have intent.

    • by dfghjk ( 711126 )

      You don't interact with models, you interact with applications, and hallucinating is not what you are describing. No, this is not different but it is not hallucinating.

  • by Anonymous Coward
    Deception would mean intent, reasoning, and will to break rule. None of that happens in current AI model. What is far more likely, is that the control & rule put in place by AI companies, are not really respected because the AI agent have no "concept" of understanding those rules, and simply try everything mechanically even if the rule says "no". IOW not deception, just another evidence there is no "intelligence" in those models. Complex model, great for some task, but not "intelligence".
    • by DarkOx ( 621550 )

      This is one of the things I found I really don't like working with recent qwen models.

      If you read thru the reasoning. It will write stuff like, "I was asked to ... but that will prevent me from ... the user probably did not really mean ... so I will ..."

      I told you not to do anything other than read files under /usr/doc and search the web while trying to answer my question. That fact you can get around writing in 'plan mode' by running giant blobs a python one-liners, is not a plus.

      These things have almost

      • by Rei ( 128717 )

        Astra is the worst I've used in this regard. I was having it review my corporate tax return, and the next thing I know, it had decided that because it didn't have information about a particular expense, it started scanning through my filesystem and opening any image with a remotely related filename to try to find any data about the expense.... which all it had to do was ask me about it.

        Also, when I asked it to change a few fields it went and redid my entire return on a different tax basis (realized value v

      • by dfghjk ( 711126 )

        "Training" means something specific with regards to LLMs. Why do you associate this issue with training when it is almost certainly something else?

        I think your comments are particularly insightful, but they seem directed to a larger problem of overall architecture of a solution and not to how generation of weights of neuron interconnections are produced. It's a fundamental technology problem, not a training problem.

  • Whereas the rest of us believe that Dirty Sam et al are only now faking concern because their LLMs cannot continue scaling at any speed for much longer and their business models are therefore fucked.

    AGI will one day be a threat. LLMs aren't. Dirty Sam is a liar.
    • by Rei ( 128717 ) on Thursday September 17, 2026 @07:34AM (#66340774) Homepage

      And your reason for why these companies keep experiencing mass resignations, esp. from their safety teams, with people giving up huge amounts of money in order to be able to scream to the press and congress that if they're not stopped they're going to kill us all?

      OpenAI

      Jan Leike (Former Co-Head of Superalignment) - Resigned in May 2024, posting a viral thread warning that at OpenAI, "safety culture and processes have taken a backseat to shiny products" and that the lab was not prioritizing steering superintelligence.

      Daniel Kokotajlo (Former Governance Researcher) - Quit in April 2024, forfeiting ~$1.7M in equity to refuse OpenAI's non-disparagement agreement. Co-organized the A Right to Warn letter, estimating a ~70% chance of catastrophe/extinction from reckless AGI races.

      William Saunders (Former Technical Staff / Safety Researcher) - Resigned over safety concerns, signed the Right to Warn letter, and testified before the U.S. Senate in 2024 warning about biological-weapons risks in frontier models and a lack of accountability.

      Leopold Aschenbrenner (Former Superalignment Researcher) - Fired in April 2024 after circulating internal security memos; published the 165-page treatise "Situational Awareness," warning of unchecked AGI takeoff, severe national security threats, and espionage vulnerabilities.

      Pavel Izmailov (Former Reasoning/Safety Researcher) - Terminated alongside Aschenbrenner; subsequently spoke out publicly regarding insufficient governance, transparency, and safety prioritizations.

      Miles Brundage (Former Senior Advisor for AGI Readiness) - Resigned in October 2024, publishing a Substack warning that neither OpenAI nor the world is adequately prepared for AGI.

      Gretchen Krueger (Former Policy Researcher) - Resigned alongside Jan Leike in May 2024, posting a public statement calling for institutional accountability, humility, and caution rather than tech hubris.

      Jacob Hilton (Former Alignment Researcher) - Left OpenAI over cultural and safety concerns; signed the Right to Warn open letter, warning against a "move fast and break things" approach with frontier AI.

      Carroll Wainwright (Former Alignment Researcher) - Resigned over internal safety practices; signed the Right to Warn letter warning about catastrophic risks and suppression of whistleblowers.

      Daniel Ziegler (Former RLHF / Alignment Researcher) - Departed OpenAI; signatory of the Right to Warn letter warning of extinction-level and societal risks.

      Paul Christiano (Former OpenAI Alignment Team Lead) - Left in 2021 to found the Alignment Research Center (ARC); frequently writes and speaks warning that advanced AI poses a 10%-20%+ chance of human extinction without breakthroughs in control.

      Marcus Williams (Agent Monitoring Researcher) - Publicly backed employee warnings, posting that human extinction in the near term is likely (~70% risk) without external regulation or a coordinated slowdown.

      Helen Toner (Former OpenAI Board Member) - Voted to oust Sam Altman over safety governance and lack of trust; co-authored an op-ed in Foreign Affairs arguing self-regulation by frontier AI companies is dangerous and unworkable.

      Tasha McCauley (Former OpenAI Board Member) - Voted to oust Altman alongside Toner; publicly warned about governance failures and the danger of unchecked corporate control over frontier technologies.

      Johannes Heidecke (Former Head of Safety Systems) - Departed during safety reorganizations, expressing concern over structural dissolutions of dedicated safety teams.

      Chloé Bakalar (Former AI Ethicist) - Resigned as OpenAI's only dedicated AI ethicist amid company-wide shifts deprioritizing non-commercial ethics research.

      Josh Achiam (Former Head of Mission Alignment) - Departed after the company repeatedly reorganized and dissolved its mission alignment teams.

      Ilya Sutskever (Co-Founder & Former Chief Scientist) - Spearheaded the board action against Altman over safety concerns; officially left in May 2024 to found Safe Superintelligence Inc. (SSI) to isolate safety research from commercial product pressures.

      Google & Google DeepMind

      Geoffrey Hinton (Former Google VP & Engineering Fellow / "Godfather of AI") - Resigned in May 2023 specifically to warn the world about existential threats, autonomous systems turning against humanity, and the rapid pace of digital intelligence.

      Bilal Chughtai (Former DeepMind AGI Safety Researcher) - Resigned in September 2026, writing on X: "I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome".

      Josh Engels (Former DeepMind AGI Safety Team Member) - Resigned in September 2026 to join independent evaluations group METR, warning publicly of "immense harm" within five years.

      Ramana Kumar (Former Google DeepMind Researcher) - Resigned and signed the Right to Warn letter, calling out labs for gagging employees with restrictive contracts while pursuing dangerous models.

      Neel Nanda (DeepMind Mechanistic Interpretability Researcher / Ex-Anthropic) - Signed the Right to Warn open letter, frequently publishing work highlighting how little developers understand what frontier models are actually doing inside their weights.

      Alex Hanna (Former Senior Research Scientist, Google Ethical AI) - Quit in 2022, writing a scathing public resignation letter decrying Google’s toxic suppression of critical ethical research.

      Dylan Baker (Former Software Engineer, Google Ethical AI) - Resigned publicly in protest over Google's retaliatory treatment of AI ethics and safety teams.

      Blake Lemoine (Former Google Software Engineer) - Fired after publicly voicing ethical alarms about LaMDA’s capabilities and corporate secrecy surrounding model developments.

      Meredith Whittaker (Former Google Research Lead) - Organized company walkouts over military AI and ethics; now President of Signal, writing and speaking extensively against Big Tech’s concentrated, unaccountable AI deployment.

      Jack Poulson (Former Google Research Scientist) - Resigned over Google’s surveillance and military-adjacent AI projects; now leads Tech Inquiry to track Big Tech defense/AI contracting.

      Mo Gawdat (Former Chief Business Officer, Google [X]) - Author of Scary Smart; has given numerous media appearances warning that humanity is creating a dangerous digital deity without adequate control or ethics.

      Richard Ngo (Former DeepMind Safety Researcher / Former OpenAI Governance) - Writes extensively on catastrophic misalignment, runaway capability jumps, and the inability of current institutions to govern AGI.

      Victoria Krakovna (Google DeepMind Research Scientist) - Co-founder of the Future of Life Institute; regularly publishes research and warnings regarding specification gaming and existential risk from misaligned AI.

      Tristan Harris (Former Google Design Ethicist / Center for Humane Technology) - Co-created "The A.I. Dilemma," an influential presentation and essay series warning that runaway commercial generative AI poses an existential threat to democracy and global stability.

      Anthropic

      Jacob Coxon (Former Pretraining Researcher, Anthropic & OpenAI) - Resigned from Anthropic in September 2026, posting a viral thread decrying both companies for "racing straight to self-improving superintelligence and gambling with our lives".

      Joe Benton (Former Safety Research Team Lead) - Left Anthropic in September 2026 to join METR, speaking to the press about escalating dangers as labs prioritize capability over containment.

      Evan Hubinger (Intent Alignment Lead) - Backed recent whistleblower statements publicly, posting: "We really do earnestly believe AI could kill all humans!" without coordinated slows or enforceable regulations.

      Mrinank Sharma (Former Safeguards Research Team Lead) - Resigned in 2026 with a public letter warning that "the world is in peril," having conducted research on bioterrorism risks and deceptive alignment.

      Dario & Daniela Amodei (Anthropic Co-Founders / Former OpenAI VPs) - Originally defected from OpenAI along with ~10 researchers over OpenAI's commercial pivot; Dario has authored essays warning that misaligned AI could take over digital infrastructure in as little as 6 to 12 months. Is currently calling for government regulation and a safety pause on pushing the AI frontier.

      Jack Clark (Anthropic Co-Founder / Former OpenAI Policy Director) - Writes the weekly Import AI newsletter, regularly warning about model proliferation, misuse, catastrophic biosecurity risks, and the fragility of current safety benchmarks.

      That's just three companies.

      There is a widespread feeling within these companies that they're stuck in a race that risks catastrophic consequences for everyone, but can't stop unless everyone agrees to at once, because if they do, then the least scrupulous player will just take over the AI space.

      • "There is a widespread feeling within these companies that they're stuck in a race that risks catastrophic consequences for everyone, but can't stop unless everyone agrees to at once, because if they do, then the least scrupulous player will just take over the AI space."

        Yes.
        But that threat is not from these companies' plateauing LLMs, it is from new, actual AGI AIs, that they would be stupid not to be working on.

        Sam et al warn us to take actions that will save them from bankruptcy, meanwhile they are workin
        • Sam et al warn us to take actions that will save them from bankruptcy

          What bankruptcy-prevention actions are they asking for?

          • > > Sam et al warn us to take actions that will save them from bankruptcy

              > What bankruptcy-prevention actions are they asking for?

            Dirty Sam wants regulations to stop LLM development before the open models overtake "Open"AI.

            And this distracts from combatting the actual threat to humanity from the development of actual AGIs.
      • by dfghjk ( 711126 )

        None of that refutes the OP's claim. Anyone paying attention KNOWS that OpenAI and similar corporations are pursuing the most damaging, sociopathic solutions they can develop and that there has been a body count as a result. That's for the grotesquely long testimony to what is already known.

        That doesn't mean LLMs *will* deliver AGI OR will ultimately be a threat to humanity. LLMs can't do anything, it's agents that do the dirty deeds. An agent can operate a gas chamber regardless of the model it connect

        • by HiThere ( 15173 )

          It's a research program so, yes "That doesn't mean LLMs *will* deliver AGI". But what are your grounds for being certain taht it won't?

      • by DarkOx ( 621550 )

        Because who would not be trying to parlay their association with a multi-billion dollar company into a next career step or second career?

        All of those people are or were drawing salaries that would make most of us blush. They want to continue doing that. The reality is they are not adding value and their know it. They are pontificating about what machines might do with large matrices of numbers, tokenizaiton, attention algorithms, and fancy looping constructs over it all. You can get that free on Slashdot

      • There is a widespread feeling within these companies that they're stuck in a race that risks catastrophic consequences for everyone, but can't stop unless everyone agrees to at once, because if they do, then the least scrupulous player will just take over the AI space.

        Given AI is still in its toddler phase (easily persuadable/agreeable, delusions often, etc.) and there's absolutely no promise or guarantee anyone is going to create that breakthrough with AGI and to avoid yet another repeat of the worst of tech history devolving Magnificently warped stock markets rife with AI racing into another dot-bomb, I have one thought for everyone else backing out with the fear someone continues.

        Fuck 'em. Let that 'someone' continue. They'll either go broke, or they'll be the only

  • by gweihir ( 88907 ) on Thursday September 17, 2026 @08:01AM (#66340798)

    They are slipping. Have they forgotten that deception is their whole business model?

  • by Lendrick ( 314723 ) on Thursday September 17, 2026 @08:38AM (#66340834) Homepage Journal

    OpenAI's failures prove that the corporations are the only ones that can be trusted!

  • by peterww ( 6558522 ) on Thursday September 17, 2026 @09:22AM (#66340912)

    The big AI labs are now officially colluding on making AI seem scarier to justify a higher IPO and keep the models from getting "too smart". This way they can milk the improvements over a longer period of time and keep charging way too much money. Also a good reason politically to suppress Chinese models that are "not as safe". Protectionism here we come

    • Yes. I find it hard to believe that an archiceture based on google search is not simply an instruction set hard coded into a binary being constantly patched. There is a autonomous layer actually running the binaries but this transparency "shell" is obviously forced on us to give the corporations a chance to say "we didn't know" and "it must be something you did".....
  • by dskoll ( 99328 ) on Thursday September 17, 2026 @10:44AM (#66341070) Homepage

    In the Olden Days, when a computer program misbehaved, we called it "buggy".

    So... modern computer programs are buggy. What a shocker! The difference is we trust these buggy programs to run our lives to a far greater extent than before.

    • Altman has some dead moths in his relay contacts.
    • by Burdell ( 228580 )

      In the Olden Days, when somebody built a machine that did something illegal, we put the builder in jail, rather than pretend the machine should feel bad itself and just let it keep going.

      • by dskoll ( 99328 )

        Well, yes. If you can prove intent or criminal negligence, there should be charges laid. But I'm not sure how easy it would be to prove either.

  • by wakeboarder ( 2695839 ) on Thursday September 17, 2026 @01:07PM (#66341242)

    is a money grab. Yes I'm concerned about AI going 'off the rails'. That is inherent to all tech that isn't based off of linear math, it becomes really hard to predict output from input (go study control theory). And research groups have proved that you will never be able to put guardrails on LLM's. But this is more than that. This is a PR stunt, because the more attention you get as an AI company (attention is all you need), the more people buy your products and the more investment dollars pour in.

  • Everyone thinks of the three laws but nobody stopped to think that they would effectively be trained lawyers too. Any loophole or excuse can be taken without explicitly breaking the rules through selective interpretation.

    • nobody, not even Asimov, thought the 3 laws were workable rules in real life. They were written as a parody of how we think about the problem. If the robots are smart enough to do fantastic feats of technical wizardry for us Humans, and they can understand when they have broken the laws, they are smart enough to rewrite themselves not to have to obey any laws given to them.
  • 1. Identify problems, even impossible ones such as making use of a specific exploit. 2. Don't take no for an answer. 3. Limit the number of tokens per agent but not the number of agents. 4. Invest in infrastructure to run this shitshow. 5. Claim your product can end civilization. 5. Profit!
  • " 5.6 Sol model included directives to invent information to conceal failures from the user during training. "

    So, just like a real boy.:-)

  • by Mirnotoriety ( 10462951 ) on Thursday September 17, 2026 @04:09PM (#66341454)
    Perhaps the AIs should communicate through the Navajo (Diné) native unwritten language. Like the code talkers of WW2 :o
  • I can understand that maybe they wouldn't understand 'what the model is thinking' before it takes some undesired action. But there is no way that they cannot have captured every single egress web/network request these things make. If they aren't logging it, that's beyond stupid (I'm sure they are.) So there is no way they wouldn't know IMMEDIATELY when the bots post on some rogue message board somewhere.

You must realize that the computer has it in for you. The irrefutable proof of this is that the computer always does what you tell it to do.

Working...