OpenAI fires multiple contractors for using AI tools to review ChatGPT responses.

Project Lily contractors are recruited through Crossing Hurdles and paid via Mercor, with some earning more than $50 per hour, according to internal documents cited by 404 Media.
In the review process, contractors summarize what users wanted and score multiple ChatGPT responses on a 1-to-7 scale, while being instructed to favor a restrained, professional tone over excessive flattery or human-like behavior.
OpenAI’s contractor operations may extend beyond Project Lily: one internal document said the company’s various projects could involve more than 10,000 contractors.
A contractor told 404 Media that AI use among reviewers is seen “all the time,” and that in a group of thousands, “tons” of people have been caught and dismissed for it.
OpenAI has fired multiple contractors who violated company rules by using AI tools—including ChatGPT itself—to review and grade ChatGPT responses. 404 Media reported that Project Lily reviewers, who earn more than $50 per hour through contractors Crossing Hurdles and Mercor, are strictly banned from using any synthetic tools because AI-generated feedback can create a dangerous AI-to-AI loop that weakens the model.
The firings expose how OpenAI's quality control efforts face internal challenges. Contractors told 404 Media that AI use among reviewers happens "all the time," and that "tons" have been caught and fired for it—raising questions about oversight across OpenAI's contractor network, which may include more than 10,000 workers.
OpenAI forbids reviewers from using ChatGPT, Grammarly, translation tools, or AI-detection software. The reason: feedback generated by AI systems can trap the model in what researchers call "model collapse." When AI trains on material created by other AI systems, the outputs become less reliable and degrade over time. 404 Media reported that OpenAI fears this feedback loop would corrupt ChatGPT's core ability to answer questions accurately.
Contractors in Project Lily review real user conversations and score multiple ChatGPT answers on a scale of 1 to 7. They summarize what users asked for and evaluate responses for accuracy, excessive flattery, and overly human-like behavior. The goal is a restrained, professional tone. 404 Media found internal documents showing some reviewers earn more than $50 per hour for this work.
Project Lily contractors review actual user conversations that may contain sensitive personal information. OpenAI says the data passes through a privacy filter and that users can opt out of training through data controls. However, the structure still exposes contractors to private details about thousands of users. The scope of this exposure is unclear, but internal documents suggest OpenAI's various contractor projects may involve more than 10,000 people.
A contractor told 404 Media that using AI to review AI "happens all the time" in the reviewer community. Among a group of thousands, the contractor said, "tons" have been caught and dismissed for breaking the rule. This suggests OpenAI's enforcement is reactive rather than preventive, and that the temptation to use AI shortcuts may be strong enough that some reviewers accept the risk of being fired.
Publishers
17
Articles
12
Reach
29