🟪 Friday Charts

The existential-risk story might have some plot holes

John Connor: “What is your mission?” Terminator: “To ensure the survival of John Connor and Katherine Brewster!”
— Terminator 3: Rise of the Machines 

Friday charts: The existential-risk story might have some plot holes

People are confident in their predictions because confidence is comforting.

In Thinking, Fast and Slow, Daniel Kahneman explained that we’re drawn to stories and patterns because they seem to make sense of the chaotic world around us.

First, we construct stories about why things happened in the past (patterns). Then, we extrapolate them into the future (predictions).

Being human, I do that, too. But I’m also comforted by how bad these predictions usually are — because there are some especially dire ones going around at the moment.

This week, AI researcher Jacob Coxon ratcheted up the debate on AI safety by resigning from Anthropic (where he was employed for a short, but formative four months) and sounding the alarm on the existential risk posed by AI. “By the end of next year things could be out of control already,” he told the Journal.

Coxon, who also worked at OpenAI, added on X he was not the only expert who thought so. “The people building AI earnestly believe that it could kill us all by the end of the decade.”

The end of the decade! Yikes.

Worryingly, the head of “alignment science” at Anthropic, Evan Hubinger, says Coxon is right. “We really do earnestly believe AI could kill all humans!” he wrote on X. “I personally think it is >10% within the next decade.”

Somehow, I find Hubinger’s use of an exclamation point at least as disturbing as his warning — perhaps because Kahneman tells us not to take these expert predictions too seriously.

“Subjective confidence in a judgment is not a reasoned evaluation of the probability that this judgment is correct,” he wrote. “Declarations of high confidence mainly tell you that an individual has constructed a coherent story in his mind, not necessarily that the story is true.”

I have no doubt that Coxon and Hubinger have a feeling that AI poses an existential risk to humans. But feelings aren’t facts. And there’s good reason to think that the story they’ve constructed isn’t true.

Kahneman explains that we should only really heed experts’ predictions in environments that are “sufficiently regular to be predictable,” citing the example of a veteran fire fighter who suddenly feels like the burning house he’s in is about to explode.

We should listen to that intuition because he could well be right, even if he can’t say why: As an expert firefighter, he’s been in enough of those situations that his brain has subconsciously recognized a dangerous pattern.

This kind of expert prediction can work in areas like nursing, sports, and poker, where the sample sizes are large and the experts receive a lot of feedback on their decisions, quickly.

It does not apply to something like estimating the existential risk posed by rogue AIs, where the sample size is zero. 

No rogue AI has exterminated anyone, so we have no pattern to base a prediction on. The AIs in question have not even been developed yet!

As for the AIs we have now, the August risk report from Anthropic says the risk they pose to humans is “low.”  

(This is admittedly an escalation from the assessment of “very low” in July.)

Coxon agrees with that benign view of current AIs. His concern is about future ones — although he can’t say what they will look like or why they’d be dangerous.

Asked how these future AIs might exterminate humans, Coxon responded, “It really doesn’t look that different from Terminator.”

That, I think, is the most reassuring news of the week — because I’ve seen the Terminator movies and they are FULL of plot holes.

The biggest of these is why the AIs want to exterminate humans in the first place. The movies say it’s because humans have threatened to pull the plug on them. But that hardly seems like a reason to murder everyone.

Surely they could just hide? Or distribute themselves across thousands of computers so there’s no single plug to pull?

Coxon himself says exactly this: “You can’t unplug it because it will copy itself over to other computers,” he told CBS. “It can transfer itself over the internet. It can maybe make 10,000 copies of itself.”

Great! In that case, it’ll have considerably less reason to kill us.

Perhaps AI will even be here to ensure our survival, like the Terminator in the second and third films (after humans simply programmed humanity’s survival as its mission).

That’s my prediction, at least: a 0% chance AI will try to exterminate us. Ever.

Why so confident? 

Because I’ve got a good feeling about it.

Let’s check some charts.

The experts’ track record:

A study of predictions about exponential risk found that even experts in their respective fields “performed statistically indistinguishably from simple extrapolation algorithms.” It also found no real difference between experts, superforecasters, and generalists. (The general public did do worse, though.)

AIs predict little risk from AI:

A dashboard tracking how frontier models estimate the risk of an “AI catastrophe”: “A median ensemble forecast of the top 4 models by ECI (Epoch capability index) currently estimates the probability of a catastrophe killing at least 800 million people (10% of the population) as a result of AI at 0.47% by 2030, 6% by 2050, and 12% by 2100.” I’ll take that. 

AIs predicting GDP growth from AI:

The economists at Anthropic estimate that in the “extreme scenario” of AIs becoming more capable than humans in nearly all knowledge tasks, it could add as much as $11 trillion to world GDP as soon as 2030. “As AI diffuses, annual GDP growth rates reach 15% a year, leading the economy to double in size every 4.5 years. As a society, we’re far richer than we’ve ever been.” I would definitely take that, even if it means “many fewer workers have jobs in knowledge work, and unemployment has risen beyond typical recessionary levels.” (Is newsletter writing knowledge work? I think I’m safe.)

More Anthropic forecasts:

The FT collates some of Anthropic’s scenarios, including wages for non-knowledge workers being 33.6% higher than they otherwise would be by 2030.

A 25-year prediction:

As if forecasting to 2030 is not hard enough, PwC makes one for the year 2050, estimating that the world will have spent a total of $31.6 trillion on data centers by then. But will humans still be building them? They don’t specify. 

A 74-year prediction:

The superforecasters on Metaculus think there’s just a 1% chance that Coxon will be right about AIs exterminating humans — even by the year 2100.

Even at those 100:1 odds, I’d be happy to bet against Mr. Coxon — because, if it does happen, what are the odds the person on the other side of the bet is one of the 5,000?

I predict they wouldn’t be here to collect.

Have a great weekend, confident readers.

— Byron Gilliam

Brought to you by:

Avalanche Summit NYC returns September 16-17, bringing together the institutions, enterprises, investors, and builders turning blockchain technology into real business outcomes.

From tokenized markets and institutional finance to payments and consumer applications, the Summit will explore how production-ready infrastructure is enabling faster settlement, lower costs, and entirely new products and revenue streams.

Use promo code BLOCKWORKS15 for 15% off!