The people paid to train OpenAI's models are getting fired for doing their jobs with the help of AI, according to a report by 404 Media based on internal documents and interviews with three contractors.

The irony at the center of the story is sharp. AI researchers have warned about "model collapse" — the degradation that occurs when models are increasingly trained on AI-generated text. Some of the people hired to prevent that outcome are themselves using AI-generated responses to train OpenAI's models, and the company's review apparatus is actively hunting them down. For more context on this story, see our ongoing breaking AI news.

"Pretty Much the One Thing That Will Get You Kicked Off"

One contractor told 404 Media they see workers using AI "all the time and people are let go for it all the time, it's pretty much the one thing that will get you kicked off ASAP." The person added that "in a group of thousands there are tons that have been caught."

The report builds on 404 Media's earlier revelation of Project Lily, an OpenAI effort in which hundreds of contractors read real ChatGPT users' prompts and conversations — which can include personal information — then rate and critique the model's responses, including checking that replies are not too sycophantic or prone to anthropomorphizing the chatbot.

The new reporting draws on additional internal documents and interviews with three contractors working across OpenAI projects. According to one internal document, those projects can involve more than ten thousand contractors.

The Rulebook: No GPTZero, No Grammarly, No AI

The documents reviewed by 404 Media are unambiguous about tool use. "Do not use AI detection tools, or AI yourself," one document for contractors who review other contractors' work states. "Do not use GPTZero or any other AI detection tool. They are not reliable. Reviewers may not use AI either, including Grammarly and AI translation."

The same document coaches reviewers on staying quiet about their methods: "Do not tell evaluators why you suspect AI. It is easier for them to hide if they know what you look for. Judge the overall pattern, not one clue."

All three contractors said reviewers are instructed not to use AI in their work, and two said people have been fired or offboarded for using it. 404 Media granted the contractors anonymity because they were not permitted to speak to the press.

So what gives a worker away? According to the report, reviewers are told to look for tell-tale signs of AI use: repetitive words, AI-style punctuation — including, apparently, overzealous use of the em dash — and work being finished unusually quickly. In Slack channels where contractors trade advice, one contractor said, people frequently post examples asking, "Is this AI?"

One contractor who admitted using AI while helping train OpenAI's models shared what they presented as their termination letter. It said their employer had identified problems with the "authenticity" of their work.

"I'm not a bad person or worker. I just needed a little boost and turned to AI to help me which eventually led to my downfall," the contractor told 404 Media. "I felt no joy in the work or that I was contributing to society in any way."

Mercor's Role and the AI-Training Labor Chain

Two of the contractors worked for Mercor, an AI-training company that hires the contractors who review ChatGPT-related material. A Mercor spokesperson told 404 Media in a statement that its experts are hired for their expertise and judgment, which the company described as essential to the ongoing advancement of AI, and that its contracts strictly prohibit the use of AI in the work.

The arrangement highlights the layered labor structure behind frontier AI development. Companies like OpenAI increasingly rely on specialized vendors, which in turn recruit large pools of freelance reviewers — by the thousands per project — to grade model outputs, flag harmful content and keep responses grounded. Quality control in that chain is enforced not by software but by other humans reading for "AI-style" tells.

Why the Ban Exists — and Why It's Hard to Enforce

The rationale for the prohibition is straightforward: if the humans rating ChatGPT's answers outsource their judgments to ChatGPT itself, the feedback loop collapses. Ratings stop reflecting genuine human preferences, model collapse accelerates, and the evaluation data that guides training becomes circular. OpenAI's own documents reportedly acknowledge that AI detection tools are unreliable, which is why reviewers are told to rely on pattern judgment rather than software.

But enforcing a no-AI rule across a workforce of thousands of freelancers, hired through intermediaries and paid per task, creates its own pressures. Contractors told 404 Media that AI use is common even as terminations are frequent, and the anonymous worker who shared a termination letter described turning to AI not out of laziness but as "a little boost" for work they found joyless.

The report lands amid a broader debate about the working conditions of the AI industry's invisible workforce — the annotators, raters and reviewers whose labor underpins every major model. It also raises uncomfortable questions for the broader economy: if the people paid to distinguish human writing from machine writing struggle to do so reliably, and AI detectors remain unreliable by OpenAI's own assessment, employers everywhere are left policing a boundary that is increasingly difficult to see.

For now, the policy inside OpenAI's training pipeline is clear, and so is the penalty. Workers grading the outputs of one of the world's most capable AI systems are expected to be entirely, verifiably human — and those who reach for the tool they are grading risk losing the job.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →