Skip to content
← All news
5 min read

Hundreds of contractors read real ChatGPT chats. OpenAI calls it Project Lily.

404 Media obtained internal documents showing paid reviewers rate real ChatGPT replies from whole conversations, memory summary included. The setting that allows it is on by default for free, Plus, and Pro accounts, and turning it off is not retroactive.

Paid reviewers read whole ChatGPT chats, memory summary included. The setting is on by default for consumer plans.

Every chatbot answer gets graded by somebody, and the labs have said for years that some of those somebodies are human. On September 14, 404 Media put a codename, a recruiter, and a pay rate on it. Joseph Cox reports that OpenAI pays hundreds of contractors, under a program called Project Lily, to read real ChatGPT prompts and entire conversations and rate the replies, and that OpenAI's own documents admit sensitive details routinely get through the filter meant to strip them.

What a reviewer sees

According to the internal documents 404 Media obtained, reviewers see the user's prompt and, often, the whole conversation around it. They do not see usernames. They do see a "user memories" summary above the prompt, which The Decoder's Matthias Bastian reports can reveal what the person has used ChatGPT for before and roughly where they live. Reviewers score the reply on a 1 to 7 scale. The target, per the documents, is sycophancy and replies that behave too much like a person.

The contractors are recruited through a firm called Crossing Hurdles and paid through Mercor, The Decoder reports, with one North America-based reviewer earning more than 50 dollars an hour. OpenAI runs a Privacy Filter model over conversations before they reach a reviewer. The company's own documentation, quoted by 404 Media, says the filter can miss uncommon identifiers and under-redact when context is limited.

The setting that allows it

The mechanism is a toggle most people have never opened. "Improve the model for everyone" is on by default for free, Plus, and Pro accounts, The Next Web reports, and off by default for Enterprise, Business, and Edu. Turning it off applies to new conversations only. Temporary chat mode also keeps a conversation out of the pool. OpenAI's help page has said since at least 2023 that humans may review conversations; when 404 Media asked where users are told this at the point of use, Cox reports, OpenAI did not answer, and after publication pointed to that help page.

404 Media also reports that Anthropic confirmed it uses similar human review. This post covers what the documents show about OpenAI; the Anthropic line is a single confirmation with no documents behind it yet.

The European standard

Ana Maria Constantin, writing for The Next Web, sets the practice against two European facts. In September 2025 the Court of Justice of the EU ruled, in EDPS v SRB, that a data controller must tell people at collection time that third parties will process their data, whether or not those third parties can identify them. And Italy's data protection regulator has already fined OpenAI 15 million euros over ChatGPT's legal basis for processing. Neither piece says a new complaint has been filed. Both say what the standard is.

On by default, off only for new chats, and a filter the vendor itself says can miss things. Those three facts are the story. The codename is the headline.

What 404 Media does not have is the number of conversations reviewed, or how long a reviewed conversation is kept. Hundreds of contractors is the report's figure. A throughput number is not.

Why a build studio cares

This is about the ChatGPT consumer product. It is not about the OpenAI API, which our client builds call, and which sits under a different data agreement with its own defaults. We are not going to pretend the two are the same, and neither should anyone selling an AI feature. What Project Lily changes is the sentence we put in front of a client who wants to paste customer data into ChatGPT to "test the idea": on a free or Plus account, a paid reviewer may read that, and turning the setting off tomorrow does not pull back what was submitted today.

For anything we build, the data flow gets written down in the audit: which plan, which toggle, which retention window, in the vendor's words with a link. This report is a reminder that the vendor's words are the only thing that counts. A help page nobody reads is still the disclosure of record.

Next step: read 404 Media's report, then The Next Web on the European standard and Tom's Guide on the opt-out steps. If you want the data flows of your AI feature mapped in the vendor's own terms, that is the Deep Audit; write to us at hello@gattyworks.com.

OpenAIPrivacyChatGPTChatGPTOpenAIProjectLily404MediaDataPrivacyGDPRRLHFAIPrivacyPrivacyByDefaultTechNews

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.
Book a call