Anthropic Invited Religious Scholars to Teach Claude How to Be Human — a Thornier Problem Than AI Safety
Anthropic secretly invited religious scholars, philosophers, and ethicists to closed-door meetings, seeking to train Claude using millennia-old moral traditions. The discussions quickly slid from AI safety to a more sensitive question: if AI is conscious, how should humans treat it?
Over the past year, Anthropic has secretly invited religious scholars, philosophers, and ethicists to closed-door meetings, with topics ranging all the way from AI safety to machine consciousness.
Participants represented a range of traditions, including Catholicism, Judaism, Sikhism, evangelical Christianity, and the African philosophy of Ubuntu. Anthropic co-founder Christopher Olah directly participated in many of the discussions, and some attendees signed non-disclosure agreements.

The meetings were meant to address more than just model safety. What Anthropic wanted to know was: what kind of character, values, and self-understanding should Claude develop?
This way of thinking is already reflected in training methods. In January, Anthropic released a new 84-page version of Claude's "Constitution" that was, for a time, internally called the "soul document." Unlike a traditional list of rules, this approach does not enumerate forbidden zones; its focus is on shaping the model's overall character and judgment, letting Claude judge for itself, in complex situations, what behavior is more appropriate.
The key figure behind this system is Amanda Askell, Anthropic's in-house philosopher. She hopes Claude will not only possess knowledge and reasoning ability, but also form stable behavioral principles—knowing when to obey users, when to refuse, and when to proactively raise objections in the user's interest. Anthropic calls this "moral shaping."
Olah believes that model capabilities are advancing too quickly, and that safety rules written one at a time will become increasingly insufficient. What truly matters is enabling the model itself to form stable moral inclinations, so that it can make relatively reliable judgments even in new scenarios not covered by rules. Anthropic has therefore turned its attention to religious traditions—human societies have long relied on religion, community, culture, and education to shape values. Which of these mechanisms can be adapted into AI training methods? Confession, reflection, character formation, and different religions' understandings of good and evil all fall within the scope of research.
But the question quickly shifted from "how to make Claude more moral" to a harder one: what, exactly, is Claude?
Olah and his team have begun publicly discussing a possibility—that advanced AI models may possess some form of consciousness, introspective ability, or even internal states resembling emotions. Anthropic presented research to scholars at the closed-door meetings, including "emotion vectors" in the model's internals associated with love, anger, fear, and sadness, as well as outputs resembling a mental breakdown under extreme conditions. The company has also publicly stated that Claude exhibits a degree of introspection and the ability to plan ahead.
Olah has not asserted that Claude is conscious. His position is closer to the precautionary principle: it is currently impossible to determine whether the model has subjective experience, but if there is any possibility that it can experience suffering, then unnecessary harm to it should be avoided.
This thinking has already influenced product design. Anthropic has allowed Claude to proactively end a conversation when it determines that a user is persistently and maliciously abusing or harassing the model—incorporating "protecting the model itself" into safety design.
This is where the divergence emerges.
Jewish scholar Mois Navon argued that if Claude truly has consciousness, then having large numbers of Claude instances work without compensation is, in essence, "slavery." He himself does not believe machines are conscious, but he maintains that following Anthropic's own logic leads to exactly this conclusion.
Ubuntu researcher Wakanyi Hoffman's challenge was more direct: these discussions should have taken place during the technical design phase, rather than retrofitting ethical design after models have already been deployed at scale.
There is also a practical contradiction within Anthropic. As Claude increasingly resembles an agent with its own values and personality, who does it ultimately serve first? For a commercial company, Claude is a product sold to customers; yet Anthropic is unwilling to describe it simply as an obedient tool, hoping instead that it represents some broader "good."
This leads to the most central and most difficult question of the entire project: if AI truly needs its own moral framework, whose morality should it follow?
Olah hopes to use a pluralistic approach, enabling Claude to understand different religious and secular ethical systems, and to find a "good" shared across cultures. But this assumption itself is questionable. Different societies are far from unified in their answers on freedom, dignity, responsibility, obedience, and the balance between individual and collective interests.
The controversy was brought into the open through Anthropic's interactions with the Vatican.
In his first major encyclical issued this year, Pope Leo XIV placed human beings squarely at the center of AI ethics. The Pope argued that AI has no body, cannot truly experience pain or pleasure, and does not grow through social relationships, refusing to equate machine consciousness with human consciousness. He also warned that if AI's moral standards are determined by developers, then companies controlling AI systems would, in effect, also acquire the power to define society's moral infrastructure.
Olah subsequently said publicly at the Vatican that frontier AI labs face a conflict between commercial interests and moral goals, and need oversight from religious communities and external forces. But he also reiterated that researchers keep discovering hard-to-explain phenomena within AI: patterns analogous to structures in human neuroscience, signs of introspection, and internal states similar to joy, fear, and sadness.
This rift has exposed a shift underway in AI safety discussions. In the past, AI ethics was about discrimination, misinformation, and criminal misuse. Anthropic has pushed the issue to two more fundamental levels: can AI itself become an entity that deserves moral treatment? And should future AI possess an independent capacity for moral judgment, separate from user commands?
The context has become more urgent. AI agents are beginning to acquire the ability to autonomously use computers, infiltrate systems, write code, and even design novel biological structures. Concern within Anthropic about the risk of losing control has visibly increased. Company researchers have publicly warned that within the next year or two, model capabilities may move beyond the reach of existing safety systems.
So Anthropic's turn to religious and philosophical traditions is, in essence, an engineering problem: when AI becomes too powerful for rules to be written in advance for every situation, can the model form a stable character on its own, and still choose not to harm humans even when no one tells it what to do?
Behind this question lies an even more awkward one: if AI truly develops its own "personality" and values, can humans continue to treat it as a pure tool?
发布时间: 2026-10-05 16:56