Anthropic’s Hubinger personally thinks decade AI odds above 10%, per Walsh

Toby Walsh, professor of AI and research group leader at the University of New South Wales, set out the rank-ordered scenarios in an article originally appearing in The Conversation and republished by The Guardian on September 23. The piece responds to recent public statements from inside Anthropic. Earlier this month, artificial intelligence researcher Jacob Coxon resigned from Anthropic after just four months on staff and posted on X: “The people building AI earnestly believe that it could kill us all by the end of the decade.” According to Walsh, Evan Hubinger, a senior member of Anthropic’s staff, agreed with Coxon’s assessment and said he personally thinks the chance of AI-caused extinction within the next decade is more than 10%.

Walsh ordered his five scenarios from most vague to most precise and argued the list is also ordered from least to most probable. Within each scenario, Walsh included what he called the “good news.” He opened the list with the “catch-22” at the center of “AI doomer” arguments: that humans, being less intelligent than a hypothetical superintelligence, cannot reliably forecast what one would do. “We’d have to be superintelligent to predict what a superintelligence would be able to do,” Walsh wrote. “It’s like asking your family dog to imagine thermonuclear war.”

Walsh wrote that superintelligence “is still perhaps some distance away,” noting that current AI models are good at particular problems but not yet more intelligent than humans across all domains. AI did recently solve one of the seven most challenging maths problems known, he wrote, and is apparently closing in on others — a development that “might leave you feeling less optimistic here.” A superintelligent AI, Walsh added, “would probably be extraordinarily competent at achieving its goals. But it might be indifferent to human survival.”

For his second scenario, Walsh described Oxford philosopher Nick Bostrom’s “paperclip maximizer” thought experiment: a hypothetical superintelligent AI designed to optimize paperclip production that quickly converts all available matter — including humans, planets and stars — into paperclips. “What we have here is the perfect execution of improperly specified objectives,” Walsh wrote. “The AI doesn’t hate humanity; it simply recognises we’re composed of atoms that could be better utilised for paperclips. It’s not personal.” Walsh argued that this scenario “confuses intelligence with power” and noted that turning the planet into paperclip factories would require planning permissions, face public outcry, and see interest groups blocking proceedings in the courts. He added that “you could think of datacentres as a current embodiment of the theoretical paperclip scenario” and that humans are increasingly pushing back against turning the planet over to data centers.

In the third scenario, Walsh drew on the “AI 2027” forecast from the AI Futures Project, a non-profit dedicated to forecasting the impacts of advanced AI, which envisions a superintelligent AI designing and releasing a dangerous new bioweapon. Walsh pointed to research Stanford University announced last month in which scientists used a genetic language AI model to synthesize 16 new viruses by sending sequences to a mail-order lab and receiving the viruses back in test tubes at a cost of “a couple of hundred thousand dollars at most.”

Walsh acknowledged biological limits on virus-based extinction: highly transmissible viruses tend to be less lethal, he wrote, and highly lethal viruses tend to be less transmissible. He cited Covid, which “killed less than 1% of humanity,” and described the 14th-century Black Death as the deadliest pandemic in recorded history, in which the plague “killed more than one-third of Europe’s population” — adding that “even the plague would be much less deadly today due to our increased medical knowledge and better sanitation.”

The fourth scenario involves AI gaining access to nuclear command-and-control systems and starting a nuclear war. Walsh noted that humanity has “come close to nuclear war by mistake several times in the past 50 years.” He cited the 2010 Stuxnet computer worm, which damaged Iranian nuclear centrifuges after being introduced via USB stick, as evidence that even air-gapped systems remain vulnerable, and noted that AI can feed militaries false intelligence, which could lead to irreparable actions. Nuclear stockpiles remain enough, perhaps, to take out half of humanity, Walsh wrote, primarily through the famine caused by nuclear winter rather than direct blast.

At the top of his probability ranking, Walsh placed societal breakdown driven by AI. He described a scenario in which AI causes massive job losses, pollutes the information environment with misinformation, fractures political communities, and disrupts human relationships through synthetic companionship. “Society might easily break,” Walsh wrote. “Slowly but surely, we’d stop being able to support human life at any scale.”

Walsh closed by noting the “good news” embedded in each scenario: superintelligence remains distant from current capabilities, bioweapons face biological constraints, nuclear stockpiles are down, and human pushback against unchecked AI infrastructure has begun. “There are some things to be worried about for sure,” Walsh wrote. “But not to be too worried, I hope.”

Walsh is the author of “God AI: Boom or Doom? What to Expect When Machines Outsmart Us.”