According to an investigation by 404 Media published on September 22, 2026, several contractors working on improving OpenAI's models were excluded from their projects for using artificial intelligence (AI) in their tasks. The investigation is based on internal documents and interviews with three contractors involved in different OpenAI projects, two of whom confirmed that their exclusion was linked to the use of AI. These individuals are not direct employees of OpenAI, but rather subcontractors. One of the affected companies, Mercor, stated that its contracts prohibit the use of language models for such tasks and that individuals found violating these rules are removed from projects. OpenAI has not commented on the matter. The contractors' role involves evaluating ChatGPT's responses to improve the system. For instance, in the Lily project, described in a previous 404 Media investigation, hundreds of people review user conversations—some of which may contain personal information—and assess whether the chatbot is responding appropriately. Evaluators check if the AI is overly agreeable or tries to mimic human behavior. However, the investigation does not specify whether the contractors excluded for using AI were working on this particular project. To understand the importance of this work, it's important to distinguish between a model's ability to generate coherent text and its ability to provide a helpful, accurate response. A well-written response may still fail to follow instructions or confidently assert false information. Human feedback is essential in guiding these behaviors. In 2022, during the launch of InstructGPT, OpenAI explained that it asked people to provide examples of good responses and to evaluate the model's suggestions. These evaluations were then used to refine the model through a process called reinforcement learning from human feedback (RLHF). However, the 404 Media investigation does not clarify how this process is applied in the current projects. In practice, an evaluator might be asked to compare two responses to the same question and choose the better one. If the evaluator uses a chatbot to make this choice, the company ends up with an AI-generated evaluation instead of the human one it requested. This issue goes beyond efficiency concerns—using AI to perform the task undermines the purpose of the evaluation. Detecting AI use without relying on AI detectors is also a challenge. According to the guidelines reported by 404 Media, supervisors are instructed not to use AI detectors, which are considered unreliable. Tools like GPTZero are explicitly excluded. Evaluators are also prohibited from using AI to write their own comments or even for translation. Instead, they look for signs such as repetitive language, unusual punctuation, or unusually fast task completion. However, they are not supposed to explain their suspicions to the contractors, to prevent them from finding ways to bypass the checks. The investigation highlights the irony that in the AI industry, relying on a chatbot to do your work could lead to job loss.