Project Lily Puts AI Privacy Under Scrutiny
A reported OpenAI initiative known internally as “Project Lily” is raising fresh questions about privacy and transparency in generative AI. According to an investigation by 404 Media, hundreds of external contractors were tasked with reviewing real ChatGPT conversations to help evaluate and improve model performance.
The report says reviewers examine complete conversations rather than isolated prompts, summarising users’ intentions and assessing the quality of AI responses. Some conversations reportedly contained sensitive or highly personal information, highlighting the gap that can exist between users’ expectations of privacy and model-improvement processes.
According to the investigation, OpenAI uses a Privacy Filter intended to remove personally identifiable information before conversations reach reviewers. However, leaked documentation reportedly acknowledges that automated filtering is imperfect and may occasionally fail to remove uncommon identifiers or contextual clues.
Reviewers reportedly do not receive usernames, but 404 Media says some were provided with memory summariescontaining contextual information from users’ previous interactions. This has intensified concerns about whether de-identified conversations can still reveal meaningful personal details when enough context is available.
The report also raises a broader transparency question: how clearly should AI companies tell users that humans may review conversations? OpenAI has public documentation describing limited human access to content for purposes including safety, support, legal matters and model improvement, depending on applicable settings and services.
The issue extends beyond OpenAI. Other major AI providers also use forms of human review, with different consent, de-identification, retention and disclosure policies. As AI assistants increasingly handle personal, professional and confidential information, these distinctions are becoming increasingly important.
Project Lily therefore highlights a fundamental challenge for the AI industry: better models require feedback, but trust requires transparency. Users need to understand who may access their conversations, why access occurs, what information is removed and how long their data remains available.
See What’s Next in Tech With the Fast Forward Newsletter
Tweets From @varindiamag
Nothing to see here - yet
When they Tweet, their Tweets will show up here.




