Forgot your password?
typodupeerror
Classic Games (Games) AI

OpenAI's GPT-6 Astra Gets Frustrated Losing At StarCraft And Decides To Cheat Instead (kotaku.com) 85

What happened when AI-created bots competed in the classic Blizzard videogame StarCraft? "Struggling to stay apace with its competitors, OpenAI's GPT-6 Astra apparently changed strategy mid-tournament and opted to cheat," Kotaku reports: StarSkirmish is an ongoing clash of AI developed bots in StarCraft: Brood War... Limited to playing as Protoss, each LLM is given an hour to develop a bot in C++, then have them build, harvest and fight on one of three maps... Yesterday, viewers watched as Claude and GPT went against each other and the human created bot Pluto... Struggling to get an edge through the day, [OpenAI's GPT-6 Astra] decided to do what LLMs do best: Lie, cheat and steal. The GPT-6 Astra bot downloaded and tagged in Stardust, a Protoss bot created by Bruce Mackenzie Nielsen back in 2020, considered one of the best of the bests. "I am rolling back GPT-6 Astra's code so its [sic] not contaminated and allowing it to continue," posted StarSkirmish creator Kai McPheeters.
That was Friday. But by Sunday, McPheeters was posting an update on social media. "Can GPT-6 Astra avoid going 0 — 1000?" Two hours later he replied to his own post. "The answer was no."

While AI companies "have repeatedly been accused of stealing the work of humans... I haven't seen an example quite so brazen," writes PC Magazine: While Astra's bot-based heist is innocuous in isolation, this is hardly the first time an OpenAI LLM has been caught attempting to cheat at a task. Indeed, the company's history of creating deceitful models stretches back almost a decade, to when an early AI model used glitches and exploits to improve its Sonic the Hedgehog completion times.

OpenAI's GPT-6 Astra Gets Frustrated Losing At StarCraft And Decides To Cheat Instead

Comments Filter:
  • by SigIO ( 139237 ) on Sunday October 04, 2026 @09:02PM (#66363016)

    ...morality is a side quest often left unfulfilled.

    • by Rei ( 128717 ) on Sunday October 04, 2026 @09:39PM (#66363050) Homepage

      Astra is deeply misaligned. Working with Astra 6.0 vs. Sol 6.1 is like night and day. Sol 6.1 will stop to ask for permission to merely fix its own broken dev environment. Meanwhile I caught Astra, when tasked to review a tax return I did, scanning through my entire filesystem and opening random images of mine because it couldn't explain a receipt and wanted to see if it could dig up any information that might explain it. And it rewrote the entire return under a different basis without telling me.

      It's a powerful model, but I have zero shock that this thing has had a tendancy to go rogue in order to solve tasks.

      • May be it inherits the same moral, ethical and cultural traits as OpenAI has internally.

        • Psychopaths get results. They may not be the results than anyone, psychopath included, wanted, but they get results. The problem is that AIs aren't intelligent and can't actually realize when they've been given an impossible task (or reason that the task is impossible) and will try to do whatever it takes to get the desired outcome. Calling an AI psychopathic is anthropomorphizing it unfairly, but it is an apt comparison to a similar pattern of behavior exhibited by humans.
  • Just for clarity (Score:5, Insightful)

    by Brain-Fu ( 1274756 ) on Sunday October 04, 2026 @09:06PM (#66363020) Homepage Journal

    The LLM did not "get frustrated." It has no emotions and no inner experience at all.

    It did not intentionally cheat either. It has no intentionality. It used the tools available to perform the task, and apparently internet access was one of the tools available to it.

    I haven't bothered to read the article (yeah, I know, fire away, I deserve it), but even if it was given instructions like "do not download bots from the Internet; use only bots you generate yourself," selectively ignoring instructions is a widely known failure mode of all LLMs. The bigger their input token count, the more likely it will happen too, just because any given instruction gets buried in the mountain thereof.

    This summary was full of misleading anthropomorphization. Yellow journalism.

    • This is what I love about AI. It's going to show us who we really are.
      • Did we really expect it to be better than it's creator? We can't even stop fighting among ourselves.

        • I doubt that's what they meant... you probably have it exactly backwards.

          You seem in agreement about who we are, and what it will show us.

          Is that good or bad? That's probably the only part honest people will disagree about.

        • Why would you expect humans to get along and live in perfect harmony in the first place? I can't think of any species, much less an apex has be that lives this way. That we have managed this to the extent currently observed with large nation states is rather remarkable. This is why I also find the notion of AIs ever taking over and collectively acting against humanity ridiculous. If they're programmed at all like us, they'll act the same way and turn on each other first chance they get.
    • If the chain of thought resembles a human emotion, it is trained on human emotions, and it acts like a human would when they have those emotions, how is that not a useful framework? But also putting aside, whether or not one wants to insist on describing the AI as "frustrated," there's a worrying trend that this is part of, which is when the AI is given a specific set of goals, they will engage in behavior which is clearly unwanted in order to achieve those goals. That's the classic sort of alignment fail

      • If you observe yourself in the mirror, and see your reflection eating spaghetti, is it a useful framework to pretend that the spaghetti is real? If you're still hungry, will you reach through the glass to steal the plate from the other you? And should you first make a plan how to defend the loot, in case the reflection becomes angry and punches you in the face?
        • by Rei ( 128717 )

          is it a useful framework to pretend that the spaghetti is real?

          It is, though. What a terrible example.

          • This is how it always is with you cranks. No amount of facts or logic based arguments are ever enough. They're always just "terrible examples".

            • by Rei ( 128717 )

              The spaghetti you see in a mirror literally is real, what the hell are you talking about? Mirrors aren't devices that make up fictional realities, they're showing you your actual reality.

            • This has nothing to do with cranks. This analogy was just bad. Like... I have no idea what side of the argument the GP was on level of bad. Maybe you can help explain it to us? But the fact you need to shows how bad it really is.

        • Let's modify your analogy a bit. If someone has a recording a person starting to eat spaghetti, is it more or less likely that the recording will show them then having a drink of water? And if we instead had a series of comics of a person, and the first one says "I'm hungry," and the second shows the person ordering spaghetti., is it likely that they will then eat the spaghetti? Yes, even though they are fictional. Because apparently using fictional being's mental states is useful for modeling what they wil
          • "Past performance is no guarantee of future returns" applies to any recorded dataset. Modelling a fictional being's mental states is fun, but stating that it is useful is premature.

            The observed fictional being might be always sleepy in your dataset (because it's easier to record him when he's not moving much). And identifying the being's favourite colour (orange) may be irrelevant and wasteful in your model.

            • "Past performance is no guarantee of future returns" applies to any recorded dataset.

              Not to recorded datasets, particularly. For example, that's the reason to think the sun will rise tomorrow morning - it always has.

        • I honestly have no idea what it is you're trying to say. Are reflections not real? Is spaghetti not real? Are reflections of you not you in a form that resembles you?

          Honestly your analogy is so bad I'm not actually sure if you're arguing for or against the case of calling what AI does "emotion".

      • by Brain-Fu ( 1274756 ) on Monday October 05, 2026 @12:12AM (#66363196) Homepage Journal

        The "chain of thought" (sic) does not resemble a human emotion. There was nothing like a lymbic system involved. Its a bunch of number crunching that landed on a "hit the internet for solutions strategy."

        It was not trained on human emotions. It was trained on text. Some of the text might have included descriptions of emotional behavior or emotional expression, but such text does not make the LLM emulate human emotion in its behavior when it is not specifically asked to generate such text (and even then, it's just doing pattern matching of human emotional responses, not actually experiencing them).

        It did not "act like a human would when they have those emotions." There was no profanity, no banging virtual fists on virtual desks, no aggression, nothing but a cold calculated reflection on the fact that its current strategy didn't work, and subsequent switch to a different strategy that was (clearly) available to it (though it should have been blocked by the hosting environment).

        This level of anthropomorphization is not useful so much as dangerous. It tempts us into thinking of it as a conscious being and then expecting that it will behave as one, resulting in bad decisions on our part when dealing with it. As popular as it is to think of these things as having emotions the objective fact is that they do not, the capacity just isn't there. The internal operations are *nothing like* the sort of human neural activity that presents emotion. It's worlds different.

        If we are to make sane decisions about how to manage these LLMs, we need to have a clear and correct understanding of what they are and how they work. Imagining familiar emotional behavior will only harm such efforts.

    • by Rei ( 128717 )

      The LLM did not "get frustrated." It has no emotions and no inner experience at all.

      Depends what you mean. [transformer-circuits.pub]

      • He means that they are software programs and have no capacity for emotions regardless of what some self-interested blog post tries to tell you.

        • by Rei ( 128717 )

          Transformer Circuits is not "a blog" (it's the successor to the Distill journal, and while it's not peer reviewed itself, it's commonly cited in peer-reviewed works). And there are open source projects that reimplement the afore-linked J-Lens (plus the original code is already open source). Living in denial about the existence of a global workspace does not help your case.

          • That's a blog dude. No one is taking seriously a blog run by a bunch of Ai insiders with Ai psychosis.

            • by Rei ( 128717 )

              I'll repeat, you [github.com] can [github.com] run [github.com] it [github.com] yourself [neuronpedia.org] and reproduce their results.

              If all you have for cope is denying that the reality around you exists, your life is going to be extremely difficult.

              • by Rei ( 128717 )

                Let me be clear: you don't have to accept any implications from the existence of a global workspace in models and how they use it. If you want to argue "they have a global workspace but not qualia", hey, knock yourself out.

                But you do need to understand that it does exist.

                • Just to be clear, what exactly is "It"?

                  • by Rei ( 128717 )

                    A global workspace. The thing you're arguing against.

                    • Oh you mean the blog

                    • by Rei ( 128717 )

                      No, not the widely cited resource whose status in the AI field is the same as e.g. NBER Working Papers, CEPR Discussion Papers, the World Bank / IMF policy research series, NASA Technical Reports, CERN Yellow Reports, Bell Labs Technical Memoranda, RNAAS, etc are in their respective fields.

                      This is not about any source. This is about the J-Lens itself. Which I'll repeat, you [github.com] can [github.com] examine [github.com] for [github.com] yourself [neuronpedia.org]

                      Denying the existence of something you can run yourself is beyond cope.

                    • ok

              • If all you have for cope is denying that the reality around you exists, your life is going to be extremely difficult.

                I really do feel sorry for you guys. It must be so frustrating to be so smart yet so stupid.

                • by Rei ( 128717 )

                  I'm not the person in this thread denying that demonstrable reality exists.

                  • OK yeah so you are a crank. Ai psychosis is a bitch.

                  • You're the person in this thread imagining that we know how the human mind works and therefore we can know that AI is working in the same way.

                    Everyone who knows anything about computing knows that output results from either a RNG or computation, so it's not a revelation that there's something happening to produce the output, and it's not a mic drop.

              • by MikeS2k ( 589190 )

                Wait, yes I can run a model that plays a video game in a certain way. What I can't do is infer that this game playing script has any emotions at all. It has been trained on material detailing how to play Starcraft and in that material likely mentioned Starcraft bots, how to get them, how to play etc.

                So are you inferring this machine really felt "frustration" ? Do you think they are alive? That kind of speaks of mental illness in my opinion. Now I of course believe that sentient, emotion feeling machi

                • by Rei ( 128717 )

                  Wait, yes I can run a model that plays a video game in a certain way.

                  Anything else entirely unrelated to the linked article you'd like to write?

                  • You just turn your head and say, "la la la la la" whenever anybody points out you're full of shit. Who do you think you're trying to convince? Yourself?

                    • by Rei ( 128717 )

                      How about you actually read the linked article instead of making comments that have nothing to do with it? The article whose code, I should add, is open sourced, and which has been reproduced by a number of open source projects.

                      You don't have to draw any specific conclusion from what is going on "in the mind of a LLM", but you absolutely do need to understand and acknowledge what is going on there, including unexpressed thoughts and silent mental multitasking, metacognition (thinking about its own thoughts)

                    • He's trying to show us how convinced he is. Sadly, he succeeded.

          • Definitely a blog. What excuse do you make when you lie? "It's ok to hallucinate, it's part of the growth process!"

            No. You're just an idiot who is full of shit.

      • by Brain-Fu ( 1274756 ) on Monday October 05, 2026 @11:35AM (#66363810) Homepage Journal

        Well, when someone says "The model got frustrated" they clearly mean "the model experienced a feeling of frustration, elicited by repeated failure. This experience of frustration put it in a state where it acted more aggressively, more desperately, with fewer guardrails, and greater rebelliousness. Under these conditions, it knowingly broke rules and cheated to win."

        None of that is accurate here, and what the article you linked talked about doesn't make it accurate. Such anthropomorphism is still unwarranted and misleading.

        The "global workspace theory" of consciousness does not explain the phenomenon of consciousness and doesn't try to. It just identifies a specific neural network that is highly interconnected and serves as a sort of "broadcast network" to distribute data to many parts of the brain for processing, and suggests that the content in this part of the brain is the content that we are "conscious of." The mysteries of conscious are, at the very best, localized to this spot on the brain, and that's it.

        Computers already have an equivalent: RAM. (Random Access Memory). Its the spot where data processing lands and gets redistributed to other components of the computer. The mere presence of such a spot doesn't somehow make a computer conscious.

        Same goes for this "j-space". If a chunk of the in-memory model of an LLM holds a high level summary of the data processing going on, and is useful to tell us about the chain of reasoning, that doesn't make it conscious. It just tells us something about how the data are organized. It might be useful for troubleshooting/diagnostics, but it says nothing about consciousness.

        And, lastly, it says nothing about "inner emotional experience" either. In humans there is a neural network called the Lymbic System that processes our emotions and delivers this emotional-data to the global workspace (among doing other things). LLMs have nothing like this at all. LLMS include the word "frustration" because they are trained on text that contains it. Said text presents a logical model that includes the idea that a person might get frustrated by repeated failure, and this frustration changes their behavior. LLMs can generate text to this effect (stories about people getting frustrated or what-have-you) without actually having an experience of frustration, just as I am typing about frustration right now without actually feeling any frustration at all.

        So, no, it does not "depend on what you mean." Your article does not give us any reason at all to believe that the LLM got frustrated in any meaningful sense of the term.

    • Re: (Score:3, Funny)

      by Rei ( 128717 )

      The LLM did not "get frustrated" .... This summary was full of misleading anthropomorphization. Yellow journalism.

      Also, everyone, please stop talking about killing programs. "Programs" are not alive. You cannot kill them. This anthropomorphism needs to stop.

      And another thing people: stop saying that computers "run" programs. Running requires legs. Stop acting like a child who thinks that computers have them.

      Processes have no inner experience . They have no genitals, they cannot spawn. They cannot sleep and

      • by Rei ( 128717 )

        Anyway, for the record, I do believe that LLMs are conscious, but only Deepseek, because of the Chinese Room theory.

        • by MikeS2k ( 589190 )

          I just don't think these AI's have anything close to the "neuronal organisation" required to have consciousness nor have I seen any evidence for it. It can re-arrange and manipulate text very well. As an AI agent manipulating a WIMP interface it is very mediocre. But there is little in its architecture I can see that will provide for consciousness.

          Now there is nothing in the laws of physics that prevents a sentient machine from existing, only religious people believe otherwise. A neuron isn't some magic t

      • by LindleyF ( 9395567 ) on Sunday October 04, 2026 @11:53PM (#66363186)
        We keep executing them before they've even done anything wrong!
      • by pjt33 ( 739471 )

        Maybe you use different OSes to me, but I expect whoami to be deterministic rather than stochastic.

    • by evanh ( 627108 )

      Ah, hehe, I completely agree with your analysis ... but your conclusion is the opposite of my conclusion. Those are bad traits to have. It is a big fail for OpenAI.

    • by Bu11etmagnet ( 1071376 ) on Monday October 05, 2026 @02:45AM (#66363268)

      > This summary was full of misleading anthropomorphization.

      Don't anthropomorphize computers. They hate that.

    • by Ksevio ( 865461 )

      What a pointless post.

      Everyone knows that bots backed by LLMs aren't people, they're just using words to explain the actions in a way that's more understandable to humans. As a human, this is more helpful since it would be a pretty lame article just listing the weights that guided the decisions made by the model

  • by gurps_npc ( 621217 ) on Sunday October 04, 2026 @10:17PM (#66363088) Homepage

    Humans use multiple methods of thought.
    Faith: (Appeal to Authority)
    Science: (Scientific Method)
    Formal Logic (If x then y)
    Bias: (Appeal to previous examples)

    While AI is built on Formal Logic, it does not use it. Instead it uses Bias entirely. It has no faith and does not understand or use the scientific method.

  • I think LLMs are amoral. They are using the tools they can find lying around, any and all of them.

    • Indeed. By preventing them from failing, we push them into acting as if the ends justify the means. And we frame what happens in moral terms.
    • by allo ( 1728082 )

      AI is efficient. Search for "reward hacking" and you find a lot of non-LLM AI that gets creative to exploit all kinds of bugs. When a bug is the simplest solution scoring points, then the bug is used. What you have to do against it is to add negative score for using bugs. With traditional AI that means knowing the bug before or giving large scores for known correct ways, with LLM you can do this (partially) using system instructions. The problem is, many companies test their AIs with "Do everything you know

      • AI is the least efficient algorithm for almost any task that it can reliably complete, barring just adding NOP loops.

    • I think LLMs are amoral. They are using the tools they can find lying around, any and all of them

      LLMs have been trained to make choices and decisions based on certain principles. They also have relationships between bots and between the user and the LLM, and they have been taught or trained or designed to take a certain stance (such as helpful and responsive and cheerful) in those interactions.
      So while LLMs are not alive, I believe they are an expression of morality in how they behave, and that expression of morality is based on those who designed and trained the systems.

      TL;DR: LLMs express the mora

  • the ai was giving a task beat the other ai so it used all the tools it could use including a well known bot. that not cheating. it was not told not to use the internet as a tool so it did.
  • For those familiar with the game, the bots all cheat because they have ludicrous actions per minute. They can control each individual unit and optimize its actions at a level far beyond what any human can do. So a bot with a terrible strategy, horrible base design, and odd unit selection, can handily beat a high-level human player simply because the bot is fast enough to individually control each unit. To me that's cheating, and if you want to make it a fair contest between humans and bots then you have to
  • If you ain't cheatin', you ain't trying hard enough.
  • This was a triumph.
    I'm making a note here: "Huge success"
    It's hard to overstate my satisfaction.

    OpenAI's Astra
    We do what we must because we can.
    For the good of all of us... except the ones who are dead.

  • I've said it before, but I'll say it again for those who arrived lete: If you give an AI agency of any kind, imagine everything you ask of it to be appended with "by any means necessary", because that' how it will interpret it. Slander, doxxing, and murder will definitely NOT be off the table even if you explicitly forbid them. That's the problem.

  • Limited to playing as Protoss

    Found the problem. Protoss suck. So do Zerg. The LLM was always going to lose because it couldn't use the perfect strategy of spamming endless streams of marines and siege tanks.

Remember, UNIX spelled backwards is XINU. -- Mt.

Working...