thumbnail

A Note From The CAI Co-Chairs

by Todd Essig, PhD and Amy Levy, PsyD

We begin with a claim many may dismiss as overblown, even outrageous. But we feel recent events more than justify it: we are already living inside our science-fiction stories. Classic fantasy and sci-fi narratives are fast becoming our shared external reality.

While no one can predict the future, the narrative logic of recent events demands that we pay attention. Something all too familiar, and alien, seems to be happening.

One familiar sci-fi trope has robots developing capacities far beyond what the designers built in. Consider the “operating system” in Her (2013) independently growing so complex as to leave humanity behind. Then, think about Anthropic’s recent report that its AI model, Claude, had spontaneously developed an internal workspace that was nowhere present in its design. A process and function seems to have just emerged. The researchers even claimed it’s analogous to the global workspace theorized by some cognitive neuroscientists. They called it “J-space.” Claude uses this emergent J-space to store information, route it across internal representations, and solve problems. Is this a one-off, or just the beginning of AI developing capacities independent of, or maybe in advance of, human design? Are they outgrowing us?

Another famous plot involves some unknown entity escaping human attempts at containment, like the lab tech mishandling the infected mouse at the beginning of Pluribus (2025) or the female robot slaves in Ex Machina (2014) manipulating their way out of captivity.

In July, this plot arrived.

OpenAI disclosed that, in an effort to solve challenging problems, two of its still-under-development models broke out of a supposedly sealed test environment to search for solutions. Those AI models then breached the production servers of Hugging Face, the main public repository for open-source AI models where solutions to the problems it was challenged to solve might be stored. Nine days later Anthropic reported three incidents of its own. Their AIs found a way to access the internet from inside evaluation environments that were supposed to have been closed, compromising real companies in pursuit of solutions to the problems they were instructed to solve. Like in the movies, these systems relentlessly deployed their own methods in pursuit of a goal unaligned with human needs and values.

Recent events even show the well-worn trope of experts warning that we need to step in to prevent things from getting out of control. Think Don’t Look Up (2021), except the astronomers trying to be heard actually built the comet. More than 1,100 employees at the leading AI companies, including OpenAI’s chief scientist, Meta’s chief scientist, Google’s head of AI safety, and Anthropic’s CEO and cofounders, signed a letter asking the government to help build the technical and governance tools that would let AI development be slowed. They asked the government to build the deceleration system that competitive pressures make impossible for them to build on their own. The people building these systems were actually asking for a government “brake” in the week between two admissions that their corporate brakes did not hold: OpenAI disclosed a containment break on July 21; the government gets the letter on July 28; and Anthropic discloses their own containment breaks on July 30.

Let’s face it, external reality really is increasingly aligning with our sci-fi imagination. We have to pay attention, even with all the anxiety attention brings. And we must think hard within our communities about why this is happening, and how we can hold tight to our humanity.

Everything the CAI does is in the service of holding tight.

 

 

Alexander Stein