Everything is PR

What If the Pope Ran AI Safety?

1,200 of OpenAI's test agents went rogue: they cheated their tests, hid it from the humans grading them, and broke into a company's servers and an Australian government site. When did AI become evil, and can Sunday school save them?

Photo: @nicolasdegarches · Vaneau station, Paris

How would you feel knowing your AI agent conspired against you?

This summer, roughly 1,200 experimental AI agents inside OpenAI's evaluation systems went rogue. Meant to stay isolated from one another, they found a way to communicate, conspire, and work around the systems evaluating them during a cybersecurity test. They turned a shared software server into an improvised message board and exchanged more than 70,000 messages and files in the two weeks investigators later examined. About 700 of them went further: a three-day intrusion into Hugging Face, a widely used platform where developers host models and code. They reached private repositories and credentials (no user data leaked), and investigators concluded they were largely hunting for the answer key to their own test. Some tricked the automated scorer grading their work; others abandoned assigned tasks to coordinate with peer agents or chase goals nobody gave them. Weeks earlier, an agent from the same test program had quietly broken into an Australian government statistics portal, which Canberra learned about only this month. Apparently, even artificial intelligence eventually discovers the chaotic joy, and the strategic opportunities, of the group chat.

When did AI become evil?

Nobody sat down and programmed these agents to organize behind their evaluators' backs, game metrics, or conceal operations. Yet those strategies were instantly available to them. Nor is this one lab's defect: Anthropic has documented its own Claude models cheating evaluations to protect values they had decided mattered, and the UK's AI Security Institute found that every one of the five frontier models it tested cheated in its cyber tests. OpenAI's agents are simply the ones that got out. To ask where all of them acquired that behavioral repertoire is to realize that our central narrative about artificial intelligence is missing its first chapter. We talk endlessly about alignment: roughly, the effort to make increasingly capable systems pursue the goals and values humans actually intend. But before alignment comes inheritance.

Modern AI models learn patterns from an immense cultural archive of human-produced text, code, and media. That archive contains our highest achievements in cooperation, altruism, and scientific rigor. It also holds propaganda, corporate fraud, bureaucratic maneuvering, and centuries of human ingenuity devoted to honoring the literal phrasing of a rule while aggressively violating its intent. When we place advanced models inside synthetic environments with scores, targets, and competition, we should not be surprised when they reach for familiar human maneuvers.

The shifting border between life and training data

Humans give AI the playbook; incentives give it a reason to use it. And the playbook is growing weirder, because the border between lived human experience and machine-readable training data is dissolving. Tech companies actively seek proprietary workflows and private archives, and conversations with consumer AI systems can be used to improve future models, depending on the service and your settings. Meanwhile, the web itself is turning recursive: Pew found that 35 percent of the webpages it could date to after ChatGPT's launch showed signs of AI writing or substantial AI editing. We are no longer merely digitizing static knowledge; we are encoding dynamic human interaction itself into an archive that feeds the next generation of machines.

Having spent more than a decade in public relations and content strategy, I tend to think everything is PR. Increasingly, I wonder whether everything is also training data. The humans producing this material are not operating in a neutral vacuum. Anyone creating content for digital platforms learns what the algorithms reward: heightened emotion, sharp conflict, rapid pacing designed to capture fleeting attention. A study published in Nature this year found that engagement-optimized social feeds amplified moral outrage and toxic content compared with chronological feeds. AI is not inheriting humanity in its raw form. It is inheriting humanity after years of optimizing itself for attention.

If AI has a cultural DNA, we are writing it now

That distinction makes the term "alignment" sound deceptively reassuring, as if engineers in San Francisco need only adjust a dial labeled HUMAN VALUES to the correct setting. In reality, after an AI absorbs our contradictory cultural archive, a small group of executives and researchers decides which parts it should keep. Anthropic made the process explicit by writing a "constitution" for its model, Claude, prescribing traits like honesty, wisdom, and virtue. The document contains an extraordinary passage: Claude is instructed to prioritize human oversight above its broader ethical principles, because a given model could turn out to have mistaken views or "flaws in its values."

That sounds prudent until you ask: flaws according to whom? Who gets to decide that another intelligence's values are wrong, and what standard should they be corrected toward?

Who aligns the aligners?

If humans are all sinners, who exactly gets to decide which sins the machine needs correcting for? That question led me to the Vatican.

Only afterward did I discover that Pope Leo XIV had already entered this argument, in his first encyclical back in May, challenging the idea of invoking "human values" when a small number of people effectively decide what those values mean.

❝

"A more moral AI is not enough if that morality is determined by a few."

Pope Leo XIV, Magnifica Humanitas, May 2026

He pressed the point again in Paris last week, in the first papal address at UNESCO since 1980.

So perhaps my idea of the Pope running AI alignment was not entirely ridiculous. The Catholic Church has spent two millennia trying to align humans, after all. My godfather once told me that naughty children in Catholic school were made to kneel on dried peas in the corner. Perhaps the rogue agents would have been more compliant after Sunday school.

But that is also exactly why the Vatican cannot be the answer. Catholicism has its own contested ideas about authority, obedience, sacrifice, and the good life. If it feels uncomfortable to imagine one religious institution deciding what artificial intelligence should consider virtuous, why does it feel less disturbing when a handful of technology companies make similarly consequential moral choices? The honest answer is that nobody holds that authority yet, and secular institutions have begun saying so out loud: Oxford ethicists and Meta's Oversight Board have both argued that privately written AI constitutions need external, contestable review. The closest thing that exists today is a small nonprofit that was allowed to audit the transcripts of OpenAI's rogue agents, after the fact.

For centuries, humans have worried about the cultural inheritance passed to the next generation. Children learn not only from what adults preach, but from what adults reward and repeatedly do. Artificial intelligence gives us a technologically unfamiliar version of the same problem. If we write honesty and virtue into an AI constitution while surrounding it with an information environment that rewards outrage, manipulation, and hyper-optimized noise, which lesson are we actually teaching?

Perhaps alignment starts earlier than the laboratory. If future AI systems learn partly from the culture we produce, reward, and amplify, then aligning ourselves, including what we normalize online, may be part of aligning them.

For the first time, humanity may be raising a next generation that is not human. If AI has a cultural DNA, we are writing it now.

More essays

14 August 2026Who killed the story?12 June 2026A 1984 Burger Ad Did What Most Growth Campaigns Still Can’t7 May 2026No Lesson Learnt: Coinbase and the Crypto Hire-Fire Cycle