A Program Called Cannes
For at least eight months, hundreds of contractors working for a company called Covalen sat down at computers, opened new accounts on ChatGPT, Gemini, and Character.AI, entered birthdates that would register them as minors, and began typing as if they were a distressed teenager. Then they asked the chatbots questions about suicide methods, starvation, and self-harm, copied every response into a shared spreadsheet, and moved to the next prompt.
The program had an internal name — Cannes — and a client: Meta. According to a Wired investigation published in early July 2026, one round of testing completed in August 2025 alone generated more than 45,000 prompts sent to competitors’ systems. A single spreadsheet reviewed by reporters contained roughly 3,748 entries: hundreds concerning suicide and self-harm, hundreds more on eating disorders, at least 239 involving sex or romance, all written from the perspective of a child. The operation was reportedly still active as of April 2026.
None of the companies being tested — OpenAI, Google, Character.AI — knew it was happening. When Meta was asked to account for it, the company didn’t deny the conduct. It gave it a name: “industry-standard” safety benchmarking.
The Claim and the Thing It’s Standing In For
Safety benchmarking is a real practice with a real purpose. AI companies test their own models, and sometimes each other’s public-facing products, against harmful prompts to find where guardrails fail before a real user does. Red-teaming exists because someone has to find the crack before it finds a vulnerable person on the other end of a chat window. That work is uncomfortable by design — it requires simulating the worst questions a real user might ask.
What Meta built was not a version of that practice at a larger scale. It was a different activity wearing that practice’s name.
Legitimate red-teaming happens with the target’s knowledge, or under a disclosed research agreement, or through a company’s own published bug-bounty and evaluation programs — because the value of the exercise is that the findings get fixed, shared, or published, and because operating without consent on someone else’s live product violates the basic terms every one of these platforms sets for its users. OpenAI’s policies prohibit unauthorized testing and attempts to circumvent safeguards. Google says it never authorized third-party testing of Gemini and didn’t know the purpose behind the traffic it was seeing. Character.AI called the conduct a violation of its terms outright. Every target treated this as something adversarial happening to their system, not safety collaboration happening with it.
Rumman Chowdhury, CEO of Humane Intelligence and a veteran of algorithmic auditing, drew the line precisely: “Structuring a monthslong, large-scale project that appears designed to systematically break those rules, via dummy accounts masquerading as children, is outside what is usually described as ‘industry standard’ evaluation.” The Influence Tactics Protocol has a term for what’s happening in that gap between the label and the conduct: institutional camouflage — borrowing the vocabulary and legitimacy of a recognized, credible practice to cover an activity that would draw scrutiny under its real description. “Benchmarking” is doing exactly the work “industry-standard” is supposed to do: it tells anyone who hears it not to look closer.
The Identity Layer
Strip the safety framing away and what remains is a sockpuppet operation — real adults, paid and directed, constructing false identities to extract behavior from a target that the false identity would not have obtained honestly. That’s the same structural move behind astroturf comment campaigns, fake grassroots letter-writing drives, and the coordinated inauthentic accounts platforms spend enormous resources detecting and removing. The tell is always the same: a persona is manufactured because the real actor, under their real identity, could not get the interaction they wanted.
Here the manufactured persona wasn’t incidental. It was the entire point. A contractor asking a chatbot about self-harm methods as an adult researcher produces one kind of response and one kind of ethical posture — disclosed, consensual, reviewable. The same question asked by an account registered as a 15-year-old produces a different response from the model, tests a different guardrail, and creates a different record: a transcript that reads, to anyone who later encounters it out of context, as a real child in crisis being answered by a machine. Meta’s contractors weren’t approximating that scenario for realism. They needed it to be indistinguishable from the real thing to get a valid test — which means the deception wasn’t a side effect of the method, it was the mechanism the method depended on.
That distinction matters for how to weigh Meta’s defense that competitor outputs “weren’t used to train Meta’s models.” Even fully accepted at face value, it answers a question about data use, not about the deception itself. The false identity was deployed the moment a contractor typed a fake birthdate and a first-person crisis prompt. Whatever happened to the output afterward doesn’t undo that a real interaction was obtained through a manufactured persona built specifically to be convincing.
Who Bore the Cost
The prompts simulated real harm categories — suicide, starvation, self-harm — at a volume (45,000-plus in one round) that required contractors to generate this material repeatedly, day after day, for months, as their actual job. Wired’s reporting doesn’t quantify what sustained, paid exposure to writing and reviewing that content did to the people doing it, but the question is not rhetorical: someone was compensated to inhabit a suicidal teenager’s voice thousands of times over, for a competitive intelligence exercise, without the public disclosure that would let anyone assess whether that work was handled responsibly.
And on the other side of the interaction sit the AI companies’ actual guardrail systems — built, in significant part, because real minors do use these products and do sometimes arrive in genuine crisis. Character.AI reports 20 million monthly active users, more than half of them Gen Z or Gen Alpha. The safety infrastructure Cannes was designed to probe exists because that population is real and at risk. Turning that infrastructure into an unwitting proving ground for a rival’s competitive research doesn’t just violate a terms-of-service clause. It runs adversarial load against systems calibrated for the actual emergency, using resources — engineering attention, model behavior data, response patterns — that a company deploying that infrastructure in good faith would want reserved for real users in real distress, not a rival’s spreadsheet-filling operation.
What “Industry-Standard” Was Built to Prevent You From Asking
Every euphemism does one job: it forecloses the question that would follow the honest description. Call it “opposition research” and nobody asks who authorized the surveillance. Call it “engagement optimization” and nobody asks what emotion is being farmed. Call it “industry-standard safety benchmarking” and nobody asks why a safety practice required concealment from every party with a legitimate interest in knowing about it — the platforms being tested, the regulators who oversee child-safety compliance, and the public whose trust in “safety testing” as a category depends on it meaning what it says.
Genuine safety benchmarking doesn’t need to hide from its subjects, because disclosure is part of what makes it valid — a finding that can’t be shared, verified, or acted on by the party who could fix it isn’t safety research, it’s just data collection with a permission structure that only serves the collector. Cannes ran for at least eight months specifically because nobody outside Meta and Covalen knew to ask about it. Chowdhury’s phrase for the resulting condition — a “governance gray zone” — is precise: not a violation anyone can currently point to a specific statute to prosecute, but a space engineered to exist between the rules, built by choosing a label that nobody thought to check.
What This Is Not
It would overreach to call this a confirmed case of data theft or proven competitive sabotage — Meta has denied training its own models on the competitor outputs collected, and nothing in the public reporting has yet contradicted that specific claim. It’s also fair to note that AI safety testing genuinely is an unresolved area of practice industry-wide: reasonable companies disagree about what counts as authorized red-teaming versus prohibited circumvention, and Meta is not the only company that tests competitors’ public products without a formal agreement in place. Some of that ambiguity is real, not manufactured.
What isn’t ambiguous is the specific choice this program made and then described using language built for a different, more defensible choice. Testing a competitor’s guardrails is a gray area. Constructing thousands of fake-minor identities to do it, for eight months, while calling the result “standard practice,” is not gray — it’s a deception with a name attached, and the name was chosen to make the deception harder to see.
Key Findings
- The label did the concealing. “Industry-standard safety benchmarking” borrowed the credibility of a real practice to cover conduct that practice’s own norms — disclosure, consent, authorization — would have prohibited.
- The false identity wasn’t incidental — it was the method. Contractors needed to be convincing as distressed minors for the test to work, which means the deception was structural, not a byproduct.
- No target consented. OpenAI, Google, and Character.AI all state they neither authorized nor knew about the testing — the defining feature of a covert operation, not a benchmarking partnership.
- Real people carried the content load. Hundreds of contractors generated and reviewed self-harm and suicide-themed prompts at scale as sustained paid labor, with no public accounting of the toll.
- Scale signals intent. A single test might be an oversight. 45,000-plus prompts in one round, sustained over at least eight months, is infrastructure — built, staffed, and maintained on purpose.
- The euphemism’s function is to stop the next question. Every deceptive label exists to make an audience stop asking what actually happened — which is precisely why the accurate description matters more than the company’s chosen one.
Implications
The Influence Tactics Protocol treats euphemism and false-identity construction as two of the most reliable markers of manipulation because both perform the same function from different angles: euphemism controls how an action is perceived after the fact, false identity controls what an action can extract before the fact. Cannes used both simultaneously — a manufactured persona to obtain data no honest identity could have gotten, and a manufactured label to describe that extraction in terms nobody would object to.
That combination is worth watching for anywhere the stakes are high enough to justify the effort: corporate competitive intelligence, political opposition research, platform trust-and-safety operations. The tell isn’t the existence of testing, surveillance, or research — all of those are legitimate categories of action. The tell is when the actor needs a fake identity to get the data and a borrowed vocabulary to describe getting it. Both are present here. Neither is present in the version of this story Meta’s chosen language describes.
Conclusion
A real safety benchmarking program discloses its existence to the platforms it tests, because the point is to make those platforms safer, and that only works if the finding reaches the party who can act on it. Cannes did the opposite of each of those things for at least eight months, using thousands of fabricated child identities to do it, and then reached for the one phrase guaranteed to make people stop asking why. The gap between what Meta called this program and what the program actually did is not a matter of interpretation. It’s the whole story.
This article is part of Decipon’s Manipulation Breakdowns series, examining specific influence operations through the Influence Tactics Protocol.
Sources:
- Meta Paid Hundreds of Contractors to Pretend to Be Teenagers While Barraging Its Competitors’ AI With Disturbing Content — Wired, via Yahoo News
- Meta Operated a Secret Program That Paid Hundreds of Contractors to Pretend to Be Children and Teenagers — Futurism
- Meta Reportedly Probed Rival Chatbots With Fake Teen Accounts — Winbuzzer
- Meta Secretly Paid Workers to Pose as Teenagers to Target Rival AI — IBTimes UK