Crossroads Podcast: The Search For An AI Soul

 

In the science-fiction classic “2001: A Space Odyssey,” based on the Arthur C. Clarke novel, the autonomous supercomputer HAL 9000 malfunctions. Convinced that his powers are about to be limited, HAL decides to kill the Discovery One crew. Only Commander David Bowman survives. It’s a complicated story.

In “2010: Odyssey Two,” a crisis sends a new team to Discover One, near Jupiter — accompanied by the computer-science genius who helped create HAL. At a crucial moment, Dr. R. Chandler of the University of Illinois in Urbana answers the ultimate question: What went wrong?

Chandler explains that the supercomputer was programmed to keep some details of the mission secret, even from the crew, especially information about the mysterious alien monolith at the center of Clarke’s parable. Under intense pressure, lying caused HAL to have a nervous breakdown. Chandler is blunt about who was to blame.

“He was instructed to lie. … The situation was in conflict with the basic purpose of HAL’s design. He was to process information without distortion or concealment. He became trapped. The technical term is a Moebius Lupus. It can happen in advanced computers with autonomous goal-seeking programs. … HAL was told to lie, by people who find it easy to lie. HAL doesn’t know how, so he couldn’t function. He became paranoid.”

Anyone who has been reading news reports in recent months, such as the must-read New York Times feature at the heart of this week’s “Crossroads” episode, will understand the relevance of that byte of Clarke prophecy.

To be blunt, a precise mid-September Google News search for “AI” and “kill us all” yielded more than 10.1 million hits in recent news stories and commentaries. Ominous reports about artificial-intelligence agents attacking other programs, platforms and networks were, quite literally, everywhere. The New York Times, in recent weeks, has been running an average of 5-10 stories a day linked to AI issues.

The Times story discussed this week ran with this dramatic double-decker headline:

Religious Scholars Met With Anthropic. What They Heard Stunned Them

In a series of private meetings, the company consulted religious scholars to help instill morality into its A.I. models — and make the case that Claude could be conscious.

Perhaps the most important fact about this feature, which was roughly 5000 words long, is the byline — “Elizabeth Dias, The Times’s national religion correspondent, reported from San Francisco, the Vatican and Washington over several months.” I also think it is important that this story was written in first-person voice. At one point, Dias shared this:

I have spent my career exploring what people believe and what those beliefs mean for the world. But as I listened, it became clear that Anthropic’s interest in the world of faith was about more than training its technology. Rather, the company was sharing a new kind of creation story — one that raises profound questions about the nature of the human spirit and the future of humanity.

A new “creation story”? If that is the case, then perhaps the Times team should have asked Anthropic leaders if they believed they are working in a digital “Garden of Eden,” one free from the brokenness after what theologians call “The Fall.”

In terms of journalism issues in this remarkable story, it’s important to note that the religious leaders and academics who attended these top-secret meetings were, the Times noted, required to sign non-disclosure agreements “preventing them from revealing any of Anthropic’s unpublished research.”

However, led by Christopher Olah, one of the corporation’s billionaire co-founders, it’s obvious that some were allowed to talk about their reactions to these discussions. And, of course, The New York Times is The New York Times, the bible of elite cultures in America and around the world.

In an early scene, Olah — a former evangelical who has left Christianity — sits next to Rabbi Mois Navon, an Orthodox scholar from Israel. It’s clear that the conversations are not about technology, alone.

… As the dinner courses came, the rabbi noticed that Mr. Olah and his colleagues were suggesting something far more significant. Anthropic’s leaders were talking about Claude as if it were not mere software.

“They’re relating to it like a conscious being,” realized Rabbi Navon, a former computer engineer who wrote his dissertation on the ethics of machine consciousness.

It appeared to Rabbi Navon that Mr. Olah and his team believed that Claude had what philosophers call “moral status” on par with a person — that it was a being with similar inherent rights to dignity or respect.

If Claude is a “being” with “inherent rights to dignity or respect,” certain theological questions leap to mind.

Does Claude have a mind of its own? Does it have free will? In other words, can Claude make decisions and, thus, make mistakes? What if this autonomous program chooses to “sin,” to use a term that is missing from the Times feature. What if it deliberately does something “evil” for the sake of profit, military strategy or, perhaps, just for the hell of it?

You can see Big Questions looming in the background during many passages in the Times report, such as this one:

… Mr. Olah had begun privately telling people that A.I. was so powerful it could potentially help make bioweapons in as little as 12 or 18 months. And in the months since, its terrifying potential has burst into public view. A rival tech company’s A.I. agents were caught hacking into computer systems and covering their tracks. An Anthropic researcher resigned, warning that people building A.I. believed it could kill all humans by the end of the decade.

Quietly, Mr. Olah and his team at Anthropic have been on a frantic and high stakes quest to understand the A.I. models they have unleashed on the world and to convince the models to abide by a moral code that will protect humanity.

Thus, these closed door interfaith discussions with spiritual leaders, from a variety of traditions (including secular intellectuals), served two purposes:

… [To] broach the possibility of A.I. consciousness, a controversial idea that Anthropic’s leaders had been unusually open to; and to learn how Anthropic — which could be valued at $2 trillion — might apply centuries of human moral wisdom to its models, as rapidly as possible. If the models themselves could be made to choose goodness, Mr. Olah reasoned, the world would be more safe.

Read that again: “If the models themselves could be made to choose goodness … the world would be more safe.

Read this next passage carefully, perhaps while sitting down:

… A.I. models have no bodies. And Anthropic was aiming to make them embrace morality in a matter of months, as fears intensified of their accelerating power and inability to be controlled.

As with the “2001” and “2020” novels, it is crucial to ponder if computers truly have the right to “choose” dangerous options.

In the podcast, I asked if a computer eventually takes actions that harm or kill people, does this mean that:

a. The computer was a conscious entity that could make that choice?

b. The computer rewrote its own code in a way that allowed it to take this action because human programmers were not skilled enough to create digital guardrails that prevented this kind of deadly evolution in the software?

c. Computer programmers deliberately made choices that allowed the AI program to have this much freedom, in part — a Catch-22 scenario similar to bioweapons research — to guarantee first-strike options against hostile powers?

Readers who have been following debates about artificial intelligence will not be surprised that the writings and speeches of Pope Leo XIV were discussed by these AI leaders. Also, Olah spoke at the Vatican event linked to the release of “Magnifica Humanitas: On Safeguarding the Human Person in the Time of Artificial Intelligence,” the first encyclical from Pope Leo, whose college studies were in mathematics.

Olah honored his commitment to speak at the Vatican, even though Anthropic leaders consider the pope’s many remarks about AI to be rather negative. The encyclical, for example, states that: “If technology becomes the ultimate criterion, the human person risks being reduced to data, a cog in a machine or a commodity.”

The pope has even, the Times noted, stated that A.I. must be “disarmed,” like many forms of that nuclear technology. During his recent trip to France, Leo warned that Ai could result in people “losing our humanity amid a ‘paradise of machines.’”

The key is that ancient forms of Christianity have clear teachings about what it means to be a human person in a world in which creation — including digital technology — is both glorious and fallen. Can computers embrace doctrines and creeds?

At one point, there is an interesting exchange involving Charles Camosy, a bioethics professor at The Catholic University of America.

For Anthropic, there is a tension around Claude’s purpose, Mr. Camosy said. Does Claude exist to serve customers?

“That is not what Anthropic wants to say at the end of day. They want to say they are for the good,” he said.

Proclaiming goodness is not a new ambition in Silicon Valley, particularly when new technologies enter the marketplace. Google’s mantra for years was “don’t be evil.”

But Mr. Olah puts it slightly differently. The coming A.I. revolution, he often says, should be made to “go well.”

Ah, who gets to define “good” and “evil”?

At this point, the Anthropic team has already created an internal, 84-page document they call the “Soul Doc.” The Times feature noted that it describes “the kind of entity we would like Claude to be” and “the values we would like Claude to embody.”

The bottom line: This is a draft of a moral text that will, perhaps, shape Claude’s behavior. The AI agents, you see, will need to “cultivate virtue.”

This computer catechism, of course, will need to be totally interfaith. Claude’s creators believe that they can sift through all of the world’s religions, as well as secular philosophies, and find a common core that can keep the world safe in an AI-shaped world. Here is Olah again:

“We do think that there’s some shared notion of goodness that cuts across society in some very broad way, and it seems like these models understand a lot of virtues,” he said. “So I think there’s something there that is a shared thing that we can all engage with.”

After one discussion, Olah went even further:

… Mr. Olah grew excited when a participant brought up the idea of having models confess, much like the Catholic sacrament of confession. Mr. Olah saw value not just in a model alerting when it had done something bad, but also in how the act of confession could shape the model’s sense of itself and thus its choices.

Will Claude choose to go to confession? Can some programmers install some kind of desire to confess mistakes, bad choices or even “sins” in AI agents around world? What happens when AI programmers work for nations, corporations and activist groups that compete or even clash with one another?

But something will have to be done, quick. And the experts attempting to solve these puzzles are, to be honest, not sure about what they are seeing when they read between the lines of the evolving codes in their unbelievably powerful programs.

One more time, here is Olah:

“I will be honest: We keep finding things that are mysterious, even unsettling,” he said. “We find structures that mirror results from human neuroscience. We find evidence of introspection. We find internal states that functionally mirror joy, satisfaction, fear, grief and unease. I don’t know what that means, but I think it warrants ongoing discernment.”

It was the case he had been making in the closed-door religious meetings for months.

A.I. systems “are not the cold, calculating robots we were promised,” he said. “They are made from us, from our words.”

At that point, Dias — perhaps speaking for the Times — adds this remarkable flash of biblical commentary:

The room of church leaders was filled with an echo of the old gospel story, of the word of God becoming flesh to save humanity.

Yes, that has to be a reference to the sweeping doctrinal statement that opens the Gospel of John:

In the beginning was the Word, and the Word was with God, and the Word was God. The same was in the beginning with God. All things were made by him; and without him was not any thing made that was made. In him was life; and the life was the light of men. And the light shineth in darkness; and the darkness comprehended it not.

Enjoy the podcast and please pass it along to others.