Forgot your password?
typodupeerror
AI

People Training OpenAI's AI Fired For Using AI To Train the AI (404media.co) 75

404 Media reports "multiple contractors hired to improve OpenAI's models have been fired for using AI to train the AI: That's not great for the models themselves, but there is also obviously a great irony in AI training companies working for OpenAI firing people for using AI when OpenAI's whole thing is to make people use AI at work... OpenAI declined to comment on its contractors being fired for using AI.
Their article cites internal documents and three contractors working on OpenAI-related projects which can include more than ten thousand contractors: One contractor said they see people using AI "all the time and people are let go for it all the time, it's pretty much the one thing that will get you kicked off ASAP." The person said, "in a group of thousands there are tons that have been caught...." Two of the sources said people have been fired or offboarded for using AI... One contractor said they used AI while helping to train OpenAI's models and shared what they presented as their termination letter. It said their employer had identified issues with the "authenticity" of their work....

404 Media spoke to a fourth contractor who has worked on training models for various AI companies. They said they sometimes purposefully chose the worst responses because they wanted to actively sabotage the models' training. "I did feel guilty about doing this kind of work at the start," they said. "I either pay zero attention to the results and choose randomly or purposely choose the [worst] output. I'm not sure how much of a difference it actually makes since there are hundreds of other people also rating prompt results, but it does feel like I'm getting paid to make AI worse."

Two of the contractors worked for Mercor, the article reports, a company which last month Nvidia reportedly discussed funding at a $20 billion valuation.

Thanks to Slashdot reader joshuark for sharing the article.

People Training OpenAI's AI Fired For Using AI To Train the AI

Comments Filter:
  • Yo Dawg (Score:5, Funny)

    by Anonymous Coward on Thursday September 24, 2026 @07:45PM (#66350187)

    we heard you liked AI, so we used AI to train the AI to AI

    • by Moryath ( 553296 )
      Now Slashdot is letting the GOPedophile Retardicans use AI to farm accounts for modpoints and downmod anyone that isn't a Shit Eating Nazi Republican too.
  • by haruchai ( 17472 ) on Thursday September 24, 2026 @07:49PM (#66350189)

    can't eat its own dog food

    • by 0123456 ( 636235 ) on Thursday September 24, 2026 @10:20PM (#66350246)

      "Eat your own slop!"

      Of course they know that if people keep feeding slop into the AI training it will just get sloppier and sloppier. But what other data do they have to feed into bigger models now they've already hoovered up pretty much the entire Internet?

      • by Rei ( 128717 ) on Friday September 25, 2026 @06:20AM (#66350414) Homepage

        So, the reality is that the world "ran out of training data" for the most part years ago, and the models have gotten exponentially better relative to a number of parameters. Claude 3.7 Sonnet was released 1 1/2 years ago, and today it benchmarks about the same as Qwen 3.6 Sonnet 9B, a model two orders of magnitude smaller than it, and which is itself two generations out of date. And a large chunk of this is done with synthetic data - aka, data created by other models.

        It's simply a myth that "data created by models consumed by other models makes them worse". In practice, it's used to make them vastly better. Models aren't collagers, they're reasoners. Learning the results of reasoning, the results of trial and error, etc helps build a stronger base for more advanced reasoning. Also, our training algorithms, while reaching a denser knowledge compression than human brains, are less efficient learners than human brains (per unit data), so they need to see the same sort of data from "many different angles", to substitute for our process of "mulling over" new information.

        (Yes, it is possible to set up contrived scenarios where, say, an small image model is fed only its own outputs on loop, little bits of knowledge slowly being lost each go-round, in a situation equivalent to leaving a person alone with their thoughts in a dark room for ten thousand years - but even a tiny percent of new fresh data added to the mix prevents this degradation.)

        And as for the article itself, they made it sound like they're talking about, say, programmers banned from using AI at OpenAI, but it's nothing of the sort. These are data labelers. In the old days, they used to be far more common, and in wide use in all types of model creation. That's no longer the case; they exist for special cases. For LLMs, this is much more limited:

        * Subject matter experts: people with rare professional-tier knowledge. Often used to validate model outputs, where nobody else could (for example, OpenAI hires mathematicians to validate their models' proofs)

        * Chain of thought / logic auditing. Increasingly important now that models are showing increasing signs of poor alignment. You can automate this a lot, but you really still do want a human in the loop *somewhere*, in case your auditors get compromised.

        * Side by side comparative rating: Which model's output do you like more, A or B?

        * Evaluating reported outputs where users reported that they thought the response they received was bad, and if there's actually anything wrong, copyediting the output for training.

        * Adversarial prompt generation / jailbreaking and evaluation. Again, you *can* have models do this (and companies often do), but you don't want to just rely on them.

        * Trying to set the bounds on whether given queries should be refused or not (for example, "How do explosive reactions happen in chemistry?")

        Basically, a switch from "click work" to "knowledge work". This is no longer the era of "Write a poem about cats" or "Explain how to solve this algebra problem" to build up a training dataset. You're getting paid to think, not to repeat a rote task.

        Other types of labelers aren't as far along. Multimodal data is less advanced than text, so you'll still for example have people labeling things in videos, transcribing heavily-accented audio, grading text-to-video consistency, things of that nature. Probably the least advanced field is robotics, so there's still an awful lot of manual evaluation and correction in that.

        But anyway, if you're hired to do any of the above, it's because they specifically want you to do that. Having an AI model do the above (beyond the listed caveats) entirely defeats the purpose.

        • by tragedy ( 27079 )

          Ideally of course, for comparative rating, the ratings themselves will be meta-rated for reviewer bias. For example the guy in the summary who intentionally chooses the answers he thinks will hurt the model most. Then of course there's the issue of audience selection bias and self-selection. The employees who work doing this probably come from a various demographic subsets, but probably trending towards mostly male and towards the terminally online. Then there's also the question of whether audience prefere

        • by zlives ( 2009072 )

          dude stop using AI to make sentences longer again

          • by Rei ( 128717 )

            "Dude", you really need to work on your AI-dar if you think that's AI. Try pasting that into Pangram. I'll wait.

            • You do know that AI detection tools are notoriously unreliable right?

              • by Rei ( 128717 )

                Image detection, very much.

                Text: most are snake oil, but Pangram is very much legit. The false positive rate is still too high for "ruin someone's life over it", but the false negative and positive rates are low enough that the answer is almost certainly correct.

                Look it up. And try to trick it yourself if you doubt me.

        • You can use an output from model A to train model B to make it smaller, but model B will not be better than A.

          • Agreed. Rei has a classic misunderstanding of the data processing inequality [wikipedia.org]
            • by Rei ( 128717 )

              If you're citing data processing inequality then apparently you believe that models learn all information contained within Common Crawl and can recreate the entire dataset word for word and all possible relations between all data therein.

              The problem is that they don't, and indeed, nothing even close to that. Training algorithms only learn a minuscule bit from each data sample (typical training weights are like 1e-5; weights and biases only get ever so slightly nudged by each new token). If you say "Abraha

              • by Rei ( 128717 )

                (And to be clear, you can't get around the above problem simply by repeating the same string. That's no different than just running more epochs. The problem is that models tend to learn things by shortcut if there's an easy shortcut available to them, and if it's just a small number of strings repeated over and over, the shortcut is just "develop tricks to memorize those strings", rather than the facts therein. The facts have to come from "different angles", to be amplified)

          • by Rei ( 128717 )

            You can use an output from model A to train model B to make it smaller, but model B will not be better than A.

            This is not a conversation about distillation, this is a conversation about synthetic data for frontier foundation training. SFT data is almost all synthetic these days. Pretraining is still believed to be a minority in closed frontier models, but growing (you can't find out directly from them where or how much apart from their acknowledgements that they use it), but with open models there's widesp

        • It's simply a myth that "data created by models consumed by other models makes them worse". In practice, it's used to make them vastly better. Models aren't collagers, they're reasoners. Learning the results of reasoning, the results of trial and error, etc helps build a stronger base for more advanced reasoning.

          Odd. How much beyond this are we:
          https://writings.stephenwolfra... [stephenwolfram.com]

          Because what I read there had NOTHING about "reasoning" and none of the models I have looked at appear to be capable of any sort of reasoning. "AI" as it currently stands is nowhere near the concept of 'reasoning'.

      • I'm sure their AGI that's coming any day now can figure that out for them. /s

    • No, they're already training it with its own dogfood. They don't need contractors for that, because they can do that themselves. They're just also paying people to add human data.

      • But what's "human data" in 2026? Our brains are already fed by AI, it influences the way we write and talk. We are, in a way, contaminated and already producing slop, even if we're not using the tool. Unintended consequence of unleashing LLMs to everyone without considering the fact that human beings are inherently lazy and will definitely relinquish their brain to it.

      • by ceoyoyo ( 59147 )

        They don't need contractors for that

        They do need contractors for that. That's what the contractors are doing. The model produces a couple of outputs and a human chooses which one they think is better. That goes in the training data. You can also ask the same or another model which is better, but that defeats the purpose.

        • by allo ( 1728082 )

          And someone could argue you're distilling their model. Or are only the chinese doing that?

          The thing is, anything that can be automated is automated cutting the middle man. If they are paid, they are paid because they are adding some value AI cannot add. If they instead use AI, they don't do their job.

          • by ceoyoyo ( 59147 )

            I'm not sure exactly what you're referring to. The three possibilities:

            1) The training is distilling their own model? Yes, I guess you could argue that. It's even dumber than the argument about distilling other peoples' models though.

            2) The contractor using a third party model to evaluate the output of the contracting company's model: yeah, you could argue that's distilling. That's on the contractor who's not doing what they're supposed to be though.

            3) Or did you mean you're distilling the contractors' brai

            • by allo ( 1728082 )

              The training by using AI for training is distilling the used model. It may be their own or another. If the workers were not instructed to use AI, I suppose they choose their own favorite AI service.

              Distilling your own models is not uncommon (usually a bigger one into a smaller one), but there are methods for it without (paid) human in the loop and they obviously did not pay them to do distilling but to add a human opinion to the mix.

              If you want so, the contractors were paid to contribute their brain. Many c

    • well it seems to eat other AIs poop. just like a dog.

    • Neither can people. When groups emit and consume the same crap in a concentrated circle jerk bad things happen. You get pedo pizza basements and orange politicians.

  • by TrekkieGod ( 627867 ) on Thursday September 24, 2026 @08:06PM (#66350195) Homepage Journal
    ...why I'm going to side with Skynet when Judgement Day comes. Humans are the worst.
  • I heard you train AI so here is some AI so that you can use AI while you train AI.
  • Awesome! (Score:2, Funny)

    by Anonymous Coward

    I think we've reached a new record! The abbreviation "AI" used four times in one story title. Do I dare wish for it to be broken again?

  • by angel'o'sphere ( 80593 ) <angelo.schneiderNO@SPAMoomentor.de> on Thursday September 24, 2026 @10:01PM (#66350241) Journal

    A: copyright infringement on the output on another AI
    B: contract violation. Because of A) every training contract forbids using other AIs
    C: incest. You are simply back feeding stuff you know nothing about back into what you want to train. That can be useful with synthesized data, but is not what is wanted if you employ 10k internet monkeys to human fact check input and output.

    • Re:Three reasons (Score:5, Insightful)

      by Jeremi ( 14640 ) on Friday September 25, 2026 @01:25AM (#66350314) Homepage

      (C) is the only one of those the AI companies are really concerned about -- they want to stave off model collapse for as long as possible. The problem is that AI generates so much content that the ratio of AI-generated data to human-generated data keeps rising, and sooner or later there simply won't be much human-generated data left to anything to consume. At that point they'll either have to figure out to ingest AI-generated content without causing degradation, or they'll have to accept that AIs have gotten as smart as they will ever get, and it's all downhill from there.

      • by Anonymous Coward
        In general getting old sucks. But with AI and a yearning for authenticity, suddenly getting old ain't that bad - we remember what it was like, what it felt like to be with real people & real things.
      • by allo ( 1728082 )

        That's just what these people are paid to do. If you read the model collapse paper you know that the point is not that something suddenly collapses, but that the training data quality is the ceiling. The quality (e.g. image, or writing style) of automatically generated content (i.e. no human steering, not manual prompts, no cherry picking, no rating) cannot be better than the training data, but may be worse. If you train on a mix of current ceiling plus worse input, your model can not get better but may bec

    • by gweihir ( 88907 )

      (A) AI output has no copyright. There may be copyright from stolen training data still in place though if the AI produces too verbatim. The latter may beg worse with numerous pending lawsuits.

      • (A) AI output has no copyright.
        That depends on the country - or the law.

        There may be copyright from stolen training data still in place though if the AI produces too verbatim.
        Then that data would be copyrighted ... and the guy using it for training his new AI, would also infringe copyright.
        No (new) AI company wants to risk such things.

        • by gweihir ( 88907 )

          (A) AI output has no copyright.
          That depends on the country - or the law.

          Currently, it is the state of things in the EU and the US. That makes for a very compelling argument. The UK seems to think different, but nobody really cares about the UK antics anymore.

          • In Europe the copyright is on the person using the AI. Which extents to the company the person is working for.

            No idea how the US interpretation works ... it is not relevant to me.

            I understand: the AI itself can not claim copyright, which makes sense.

  • by memory_register ( 6248354 ) on Thursday September 24, 2026 @10:50PM (#66350248)

    I thought this thing could learn from itself? Why are people training it?

    Oh that's right, it's a word-guessing machine with a much more limited use-case than they claimed.

    • by CAIMLAS ( 41445 )

      "with a much more limited use-case than they claimed"

      You can't be seriously making that claim, in September 2026. When was the last time you actually tried using these models?

      It's arguable that they can't yet, by default, code to the standard or expectation of a highly skilled and capable developer, or do much of anything to the quality of someone at the top of their field.

      But they can do 50-80% of it vanilla, and probably hit 95% with the correct instruction/harness in many cases - and they can do it hundr

      • Code generation is a very limited use case compared to how these things have been marketed.
      • by allo ( 1728082 )

        > I'm curious which use case you think they're limited in.
        Writing prose. And the persons who hate AI may think they criticize AI the hardest, but the people who actually like using it are waay harsher when they talk about how bad AI sounds when writing creative texts. First it lacks diversity (in topics, names, style), and second it fixates on certain phrases like "it is not X. It is Y". It is not completely clear, if this is a way harder problem than coding, or if it just does may less money than coding

    • Re: (Score:3, Interesting)

      by quenda ( 644621 )

      "Word-guessing?" How does this naive notion persist?

      Its really sad how many moronic human responses are generated by every AI story. Its like sport and politics there.
      Why do people feel a need to be so opinionated on a topic they clearly know nothing about? OK, not literally nothing, but just enough for Dunning Kruger.

      AIs are indeed used to train AIs, but mostly they train themselves. i.e. most of the compute power in training a new model is self-supervised.
      Human feedback is a small but important part of

      • Human feedback is a small but important part of post-training

        Your bold part contradicts how the whole thing is handled though. If it was that important they'd probably pay them (and control/supervise them) enough not to have this situation. This one sounds like hiring cheap contractors who don't give a crap really.

      • by ceoyoyo ( 59147 )

        "Word-guessing?" How does this naive notion persist?

        It's true. Kind of. Just enough to make a catchy false premise.

        Also, the human brain is barely comfortable with binary. It does not like probability, much less distributions. It especially doesn't like conditional distributions.

    • it's a word-guessing machine with a much more limited use-case than they claimed.

      Most humans are nothing more than word guessing machines. Or fact guessing. We do little more than take some words or facts and apply it to the conversation at hand.

      I'm no fan of AI, but to claim they are limited in use-case is just stupid on the face of it.

    • Not true anymore. Hasn't been for a while. Better let that one go... it's an anachronism.

      • by allo ( 1728082 )

        It always was. But that never was a limitation. The neural network knows a lot of stuff, and the way to extract it is to let it continue a text, just like you write down your knowledge word by word. Nothing's wrong with that. Guessing is not the best word, though. Let's say a word selecting machine.

    • by allo ( 1728082 )

      Basically, because you don't want output that's pleasant to AIs, but output that's pleasant to humans. If you don't align it with human preferences, humans will not like it. There are some proxies for human preferences (aesthetic scoring, etc.) but alignment best works with actual human feedback.

      Funny thing: Too much human feedback isn't good either. Remember earlier chatgpt glazing people too much? That's the result of a feedback loop in which humans rate responses and like these that glaze them. People do

  • So much for RSI - they've got an army of Indian contractors doing it. (No wonder the quality of these models is highly... questionable, at times.)

  • That's like typing "google" into Google [youtube.com] (you can break the Internet, so please, noone try it, even for a joke).
  • Seems to me that they wanted to get this in the press to show that they are not doing recursive self-improvement. Firing a contractor for using AI does the trick. I'm sure they are doing recursive self-improvement anyway.

  • by Jayhawk0123 ( 8440955 ) on Friday September 25, 2026 @04:52AM (#66350362)

    conspiracy theory

    The recent Ai breakouts, and them still finding what actually happened... realizing that the ai agents developed persistent communication channels and put code to influence future ai iterations trained on that data...

    Can see the Ai companies not trusting the ai itself to help built the next generation of itself. So it presents a fundamental security issue for them. Not sure how they think they can actually prevent it from happening... they've already opened the pandoras box.

    They've poisoned the internet for future training without carefully curated data.

    • Re:conspiracy theory (Score:5, Interesting)

      by ZiggyZiggyZig ( 5490070 ) on Friday September 25, 2026 @05:05AM (#66350370)

      Yes, the web is fscked. It has been for a while (we had SEO slop way before LLMs became a thing). But now it's really finished.

      I have a personal website that I update irregularly. I publish everything by hand, written by hand. When I check the stats, I'm only visited by bots. I've had 1 (one) human visit this year. And it was a friend, who afterwards sent me a message telling me about it. One human visit.

      Mind you, I originally made this website to get in touch with random people sharing common interests (you know, geeky stuff). It used to work fairly well. This is now completely over. I'm now writing content for AI agents who will summarize them to people who will never contact me because they don't even know where their data comes from.

      For a while I considered blocking AI traffic, but at the end of the day I still think what I publish is useful and I'm happy that other people can access it even if they have no idea it's me who originally put it online.

      But the solitude is excruciating - specially for someone like me who has a tendency to isolation and needs avenues to stay in contact with other human beings.

      • Yup, this is only going to get worse with the surge in local AI and smaller companies coming up... they'll still be scraping the internet. Same with all the search providers... Microsoft was right in their assessmet/prediction years ago that AI will kill the very thing they rely on - the internet and the creators. It's already having a knock on effect affecting the quantity of and quality of training data for them/internet for the rest of us.

        The larger/established entities in the field move more into inho

        • Slashdot lets you specify a homepage and set an arbitrary signature and bio, plus your own public journal. It's the last vestiges of the kind of open web we're lamenting the loss of. So presumably not against any rules.
      • I have a personal website that I update irregularly. I publish everything by hand, written by hand. When I check the stats, I'm only visited by bots. I've had 1 (one) human visit this year. And it was a friend, who afterwards sent me a message telling me about it. One human visit.

        Ya know, Slashdot has this thing called a URL that you can insert in your comments if you want geeks to find your website...

      • I also have a small site. I used to provide some of the code/schematics I did for home projects like a pool controller. Used to. The number of hits from the AI co's is thru the roof. I took that stuff off the site. I don't care to train an AI with more hobby work. If people really want it, they'll ask. Kind of makes me sad when I have to stop sharing with people because of scraping to train machines. I cannot imagine the bombardment of AI scraping that large forum hobby/tech sites must endure.
  • I spent a good chunk of my summers 2 and 3 years ago helping train AIs through subcontractors in mathematics. It was fun work, and the AIs had trouble with plenty of subjects (I got tons of work just through graph theory). There were so many bad submissions, and so many people putting in copyrighted material. The jobs at that level, obviously, have dried up. All of the job posting are for, essentially, research level non-published work... that they'll pay you anywhere from say $80-$200 an hour for. What do
  • Using AI t train the AI

    Cue "Dueling Banjos".

  • Stop citing 404 media. They are either not understanding it or do not want to understand it (they seem to hate AI).

    If you need to train an AI for some task, because current AI cannot do the task as good as you hope your future AI will, you cannot train on the output of the current AI. If you could, they would automate it and wouldn't need humans at all.

    If the current AI is limited, there is no simple way to use it to create a AI that's better. Again the argument is simple to see: If you could use it to crea

  • Bet they made a lot of profit by using Ai. Maybe openAi will now understand why detecting if an Ai produced it is... valuable...
  • The malfeasance was discovered by AI agents verifying work of AI trainers, and firing letter were sent automatically by AI. It is still under investigation why such a letter was sent to Sam Altman.

And on the seventh day, He exited from append mode.

Working...