
Hundreds of contractors, some paid more than $50 an hour through staffing agencies, have been reading real ChatGPT prompts under an OpenAI programme codenamed “Project Lily”. Over at Microsoft, “at least hundreds” of reviewers grading Copilot’s photo editor have been wading through upskirt shots and requests to undress real women. And Meta’s shiny Muse agent, sold as an AI that phones businesses for you, was at times a person in a call centre.
All three stories broke in the last three weeks, most of them via 404 Media. Taken together they describe an industry whose “artificial” intelligence keeps turning out to have a pulse, a timesheet and, in at least one case, a very unpleasant Tuesday afternoon.
Project Lily and the people behind the curtain
404 Media’s Joseph Cox reported on 14 September that OpenAI is hiring hundreds of contractors to read a steady stream of real user prompts. Each reviewer reads the prompt, summarises what the user wanted and rates four possible ChatGPT answers on a scale of one to seven, according to Tom’s Hardware and Android Headlines. The point, charmingly, is to make ChatGPT sound less robotic. Nothing makes a chatbot feel human like a human.
Reviewers are recruited through a firm called Crossing Hurdles and paid through Mercor, the AI-training labour broker. They do not see usernames, and OpenAI runs a “Privacy Filter” model to strip personal details before a prompt reaches them. OpenAI itself concedes sensitive details can still slip through. Reviewers can also see a “user memories summary”, the running profile ChatGPT keeps on you, which can reveal your job, your rough location and your personal history.
ChatGPT is used by more than 900 million people, many of them as a therapist, a doctor, or a 2am confessional. One contractor, asked by 404 Media whether users realise a stranger might be reading, put it bluntly: “I don’t think they would imagine some contractor somewhere is analyzing the conversations.”
When asked directly whether users are told humans read their chats, OpenAI initially had no answer. It eventually sent a link to an FAQ page that mentions human review for model improvement, per IBTimes UK. So the disclosure for 900 million people is a help page, which is plenty, obviously.
Copilot’s reviewers got the worst of it
On 28 September, 404 Media followed up with Microsoft. Contractors hired through platforms including Prolific grade Copilot’s image-editing feature: they see the user’s original photo, the prompt, and two AI edits, then pick the better one. Faces are reportedly left unblurred, per Digital Trends.
What they got, according to TechSpot and Malwarebytes, was a firehose of requests to enlarge real women’s breasts, shorten their skirts or pose them sexually, alongside upskirt photos, pro-anorexia material and foot-fetish images of children’s cartoon characters. “I just came across an image set that consisted of eight upskirt photos,” one reviewer wrote. Another said simply: “I recoiled.”
The grim detail is the job description. These workers are there to judge quality, which in practice can mean deciding which of two sexualised edits of a stranger better matches what the user asked for. One contractor asked, fairly: “Who is writing these prompts and who is deciding that basically generating porn is what Copilot is now focused on?”
Microsoft’s full response: “Microsoft uses customer data as described in our terms of use, including to improve our products and enforce our code of conduct.” Twenty-two words of legal wallpaper, pasted over a workforce being fed other people’s abuse material by the hour.
Meta’s AI caller had a human voice
Then there’s Muse, Meta’s new assistant, which can phone US businesses to book a haircut or haggle over a bill. On 22 September, 404 Media and Reuters reported that Meta had added “a human agent layer for calls to get completed” during internal testing, because businesses kept hanging up on the bot. With humans in the loop, success rates hit 95% to 98% in some tests.
Testers were not always told a person had made the call. One Meta employee warned internally that it “could portray us as ‘their AI is not good enough so they still need humans'”. (It could indeed, mainly because they did.)
It got worse. An employee who asked Muse to negotiate his cable bill found a transcript showing the contractor had made a racist reference on the call, per Gizmodo. A vice president in Meta’s Superintelligence Labs apologised, called the undisclosed human calls a mistake, and said the feature had been rolled back. That’s the Superintelligence Labs, whose flagship agent needed a call centre to book a trim.
How we got here
In fairness to all three companies, human review is how these systems get built. Rating answers is the “human feedback” bit of reinforcement learning from human feedback (RLHF), and Google openly says trained reviewers read a subset of Gemini chats. None of this is secret in the legal sense; it’s buried in privacy policies nobody reads.
The pattern is old, too. Facebook’s “M” assistant in 2015 leaned heavily on human operators. Amazon’s “Just Walk Out” checkout relied on around 1,000 workers in India, with roughly 700 of every 1,000 sales in 2022 needing human review. The AI boom… err, sorry, the “agentic revolution” keeps discovering that the cheapest way to make software look clever is to hide a person inside it, like a chess automaton in a trench coat.
| Company | What humans did | Reported | Company line |
|---|---|---|---|
| OpenAI | Read and rate real ChatGPT prompts (“Project Lily”) | 14 Sept | Pointed to an FAQ page |
| Meta | Placed Muse “AI” calls from a call centre | 22 Sept | Called it a mistake, rolled it back |
| Microsoft | Graded Copilot photo edits, including sexualised ones | 28 Sept | Cited its terms of use |
What you should actually do
Settings as of 30 September 2026. In every case, opting out affects future chats only; anything already reviewed stays reviewed.
- ChatGPT: Settings, then Data controls, then switch off “Improve the model for everyone”. This is on by default for Free, Plus and Pro. For one-off sensitive questions, use Temporary Chat. Business, Enterprise and Edu accounts are excluded from training by default.
- Microsoft Copilot: on copilot.com, click your profile icon, then your name, then Privacy, and turn off “Training on conversation activity” and “Training on voice conversations”. On Windows or Mac it’s Settings, then Privacy; on mobile, Account, then Privacy. Microsoft’s own page says data may still be used for safety, ads and “other general product” improvements, and says nothing about opting your uploaded photos out of human grading.
- Google Gemini: go to Gemini Apps Activity and turn off “Keep Activity”. Google admits human reviewers read a subset of chats, and those are kept for up to three years even if you delete your history. With the setting off, chats are held for 72 hours.
- Meta AI: in the US, there’s no switch. Your Meta AI chats also feed ad targeting. In the UK and EU you can file an objection through Meta’s Privacy Centre. Otherwise, the only reliable setting is not telling Meta AI anything you’d hate a contractor to read.
- Claude: Settings, then Privacy, then turn off “Help improve Claude”. With it on, Anthropic keeps your data for up to five years; off, it falls back to 30 days.
Rule of thumb: if you wouldn’t say it to a stranger earning $50 an hour on a gig platform, don’t type it into a free chatbot.
What this means
- “Private” AI chats are private in the way a hotel room is private: housekeeping has a key.
- The reviewers are the ones taking the damage, absorbing trauma for piece-rate pay while the companies issue one-line statements.
- Disclosure is the scandal. Every firm has a policy page; none of them told you in plain English at the moment you typed.
- Opt-outs exist for four of the five big assistants. Use them today.
We covered Muse moving onto your Mac to read your files this week, and ChatGPT plugging into patient records earlier this month, so the amount of intimate data flowing through these systems is only going up. If you’d like someone reading the fine print so you don’t have to (a human, admittedly, but one who’s on your side), sign up to the TopToolStack newsletter.
Did you know: Google’s human-reviewed Gemini chats are stored disconnected from your account, so deleting your history can’t reach them for up to three years.
Sources
- 404 Media: Inside ‘Project Lily’: The Humans Reading Your ChatGPT Chats
- Tom’s Hardware: ChatGPT transcripts are reportedly read by humans
- 404 Media: Humans Are Reading Copilot Prompts
- Malwarebytes: Humans are reviewing Copilot users’ image-editing requests
- TechSpot: Copilot’s human reviewers can see users’ uploaded photos
- 404 Media: Meta Tests Muse AI Agent Calls Made by Humans in a Call Center
- Gizmodo: Meta’s Muse agents sometimes tagged in human contractors
- Microsoft Support: Copilot privacy controls
- Google: Gemini Apps Privacy Hub