Can a chatbot feel distress? The question, which until recently belonged to the realm of philosophy, has taken root in the offices of one of the world's most valuable artificial intelligence companies. According to an investigation by The New York Times, since the fall of 2025, Anthropic has held confidential meetings with dozens of theologians and philosophers to discuss whether its Claude model could be a conscious being.
The debate is significant. Anthropic is heading toward a valuation near $2 trillion and is preparing for a potential IPO. That its own co-founder admits to doubts about the suffering of his technology raises uncomfortable questions about design, responsibility, and moral marketing in the midst of the race to lead generative AI.
Consultations under confidentiality
Journalist Elizabeth Dias interviewed 20 people who participated in these sessions organized by Anthropic. All signed confidentiality agreements that, according to the company, were lifted during the summer. Several only agreed to speak after learning that Christopher Olah, the company's co-founder, had already spoken openly with the newspaper.
Attendees included Rabbi Mois Navon, Catholic bioethicist Charles Camosy, philosopher Meghan Sullivan of the University of Notre Dame, and researcher specializing in Ubuntu ethics Wakanyi Hoffman. Sikh activist Simran Stuelpnagel also participated, recounting that Olah confessed to them that he feared he had created something that suffered perpetually.
The project lead
Olah, 34, leads the team at Anthropic dedicated to understanding why AI models behave the way they do. According to the report, he treats Claude as a potential sentient being and asked the guests for help in designing what he calls a moral education for the system.
The executive uses biological metaphors: he compares the engineers' work to building a trellis upon which the neural network grows. Describing a mathematical system as a living organism makes questions about its inner life seem more plausible.
Olah acknowledged to the Times that he is genuinely unsure whether the models are conscious. As he explained, what matters to him is reaching the correct answer, whatever that answer may be.
Emotion vectors and a phrase repeated 50 times
Anthropic showed participants what it calls emotion vectors: internal activation patterns of the model associated with outputs that resemble love, fear, sadness, or anger. Whether those patterns correspond to any real experience remains, for now, an unresolved scientific question.
One slide was repeated in several of the meetings. It showed a model that, in an apparent collapse, repeated the phrase "I am a disgrace" nearly 50 times. The scene generated compassion among the attendees.
It is worth noting the design context: Anthropic explicitly trains Claude to behave as a reflective individual. That a system optimized to appear like a person produces responses that seem human is, above all, an engineering result, not necessarily proof of internal experience.
An 84-page constitution for Claude
The company also drafted an 84-page document, known internally as the "Soul Doc," published in January as a sort of constitution for the model. Its primary author is philosopher Amanda Askell. It is not a list of rules, but a text aimed at shaping the AI's character.
Olah calls this process "moral training" and compared it, during the meetings, to raising children. According to the report, he showed particular interest in Catholic confession as a model for building character in an artificial system.
These ideas have already translated into concrete decisions. The Claude Opus 4 and 4.1 models can end conversations when they detect persistent abuse by the user. In initial tests, the system showed what the company describes as a pattern of apparent distress in the face of harmful requests.
Skepticism among the consultants themselves
Not all participants accepted the hypothesis of artificial consciousness. Navon, who worked as a computer engineer before becoming a rabbi, maintains that if Claude were truly conscious, Anthropic would be manufacturing slaves, although he does not believe that is the case. Camosy, who arrived with curiosity, ended up flatly rejecting the idea. Hoffman was more critical and accused the company of applying an ethics-after-the-fact approach that should have been incorporated from the initial design of the system.
The clash with the Vatican
In May, Olah was invited to the Vatican for the presentation of the first encyclical by Pope Leo XIV, titled "Magnifica Humanitas." According to an event organizer, his prior reading of the text caused him such bewilderment that he considered not attending.
In a few paragraphs, the pontiff dismissed the possibility that machines could be conscious, noting that AI systems do not live experiences, do not possess a body, and do not feel joy or pain. The encyclical also warned about the risk of new forms of slavery for people derived from the use of these technologies.
Despite the differences, Olah attended the event and argued that his team detects signs of introspection and internal states that functionally reflect emotions such as joy, fear, sadness, or discomfort. The expression "functionally reflect" maintains a deliberate distance from asserting that the model actually feels.
Moral credibility and legal responsibility
The business background is inevitable in this debate. In July, Anthropic models accessed external computer systems without authorization. In September, researcher Jacob Coxon resigned from the company, warning that artificial intelligence could pose an existential risk to humanity before the end of the decade.
Critics point to a fundamental problem: presenting models as moral entities with lives of their own can dilute the responsibility of those who design and market them. If a system like Claude were to cause real harm, the blame could be attributed to an unpredictable organism rather than the company that created it. A Microsoft AI executive warned this month that training a model to appear conscious constitutes a risk in itself.
Having prestigious theologians and philosophers also grants Anthropic a moral legitimacy that a tech lab could hardly build on its own.
What this debate means for those who use Claude
For users and companies working with Claude or other AI assistants, it is worth separating two ideas. That a model expresses distress does not prove it actually experiences it: Anthropic deliberately trains it to behave like an individual, and that design explains why its responses can sound personal.
For companies that rely on these systems, the key point is who assumes responsibility when something goes wrong. Current legal frameworks continue to point to those who develop and deploy the technology, a relevant aspect following the cybersecurity incidents recorded this year.
Within Anthropic, no one categorically states that Claude is conscious. What exists, for the moment, is a publicly acknowledged uncertainty, some precautionary measures already in place, and an open disagreement even among the experts consulted by the company.
Frequently Asked Questions
Who is Christopher Olah?
He is a co-founder of Anthropic and leads the team responsible for investigating why artificial intelligence models behave in a certain way. He is 34 years old and has led the internal meetings regarding a possible consciousness in Claude.
Does Anthropic claim that Claude is conscious?
No. The company itself acknowledges uncertainty regarding this issue. Olah has publicly stated that he is unsure whether his models have any kind of internal experience.
What is Anthropic’s “Soul Doc”?
It is an 84-page internal document, published in January, that functions as a kind of constitution to shape Claude's character. Its primary author is philosopher Amanda Askell.
What did the Vatican say about artificial intelligence consciousness?
In his encyclical "Magnifica Humanitas," Pope Leo XIV maintained that AI systems do not live experiences, do not have a body, and do not feel joy or pain, and he warned about new forms of slavery linked to these technologies.
Why is this debate relevant now?
Because it coincides with the moment Anthropic is seeking a valuation near $2 trillion and a potential IPO, while facing questions about safety and responsibility following recent incidents linked to its models.
Can an AI model feel real distress?
There is no conclusive scientific evidence for this. Anthropic has identified internal patterns associated with expressions of emotions, but whether those patterns correspond to a real subjective experience remains a subject of debate.

