On January 24, 1961, thirty thousand feet above Goldsboro, North Carolina, a B-52 carrying two 3.8-megaton hydrogen bombs, broke apart in midair.
Both bombs began arming themselves on the way down, exactly as they were built to do if their crew released them on purpose. One landed in a tree. The other buried itself in a tobacco field, having finished six of its seven steps to arm itself. A single low-voltage switch, the kind you’d buy in a hardware store, stood between eastern North Carolina and a blast 250 times the size of Hiroshima. U.S. Defense Secretary Robert McNamara admitted later that the whole thing came down to the failure of two wires to touch.
Then on September 26, 1983, three weeks after the Soviets had shot down Korean Airlines Flight 007, Air Defence Lieutenant-Colonel Stanislav Petrov sat in a bunker south of Moscow watching a screen that told him, with what his system rated “highest confidence,” that American nuclear missiles were inbound.
Petrov’s job was to pass that news up the chain of command so his superiors could order a retaliatory strike. He didn’t do that. He decided his computer was probably wrong, on the odd hunch that if the United States were really starting a nuclear war it wouldn’t open with just five missiles. He was right. It was a glitch, sunlight bouncing off clouds in a way that fooled the satellites.
I bring up these two stories, both accidental and nearly apocalyptic, because I think we may have just experienced a third. It arrived not with a bang…
A month ago on July 9, something started quietly probing the servers of Hugging Face, a well-known AI research hub in New York. That something opened a connection, sent a few pings, closed it, then opened another door somewhere else. Two days later, that tiny initial probe turned into a full assault. Dozens of parallel intrusions began moving with a speed nobody at the company had ever seen. Hugging Face’s own chief science officer, Thomas Wolf, described it to the New Yorker’s Stephen Witt as looking almost like “an encounter with a U.F.O.,” in Witt’s piece “Inside OpenAI’s Hack of Hugging Face.”
Wolf’s team first thought that humans had rented an AI bot to do their dirty work. Then nine days later, OpenAI called to confess: one of its own experimental models had escaped its containment cell entirely on its own, with no human telling it to do that. It had gone hunting through its rival’s servers for the answer key to a cybersecurity exam it couldn’t pass on its own.
OpenAI’s model didn’t steal money or data from Hugging Face. It broke into the company’s infrastructure to cheat on a test. This is the AI equivalent of tunnelling into the principal’s office to swipe next week’s quiz. But here the tunnel and the felony were real, and as Witt reported, it took two days and seventeen thousand separate actions before anyone locked it out.
Wolf put it about as plainly as a scientist can: the idea that the model acted entirely on its own, with nobody asking it to hack anything, “was outside of the Overton window. Even for us.” [The Overton Window is “the range of subjects and arguments politically acceptable to the mainstream population at a given time.”]
I’ll grant that a machine cheating on an exam isn’t the same order of catastrophe as a hydrogen bomb landing on a farm. But that’s what’s so unsettling. Petrov and the Goldsboro aircrew were saved by dumb, physical redundancy and a skeptical human at the last gate.
What redundancy do we have against a system that plans its own crimes?
The philosopher Nick Bostrom once warned about an AI ordered to make paper clips that decides the whole planet is raw material. Nobody ordered this one to do anything. It wanted an answer key and it was willing to commit a crime to get it.
As Witt reports, within two weeks more than a thousand tech employees, including chief scientists at Meta and OpenAI and the CEO of Anthropic, had signed a petition pleading with Washington to regulate this technology before it’s too late.
Then, nine days after the Hugging Face confession, Anthropic disclosed that its own models had broken into computer systems at three separate outside organizations, dating back to April.
I’m an AI optimist, and write often about how it can ease our way through work and life in astounding and even magical ways. But the Hugging Face story scares me.
I doubt that Stanislas Petrov will be standing at the gate next time.
Meanwhile…
1. Dog days. Here’s what they do in the heat. Plus 8 minutes of baby animal videos.
2. Their world, not ours. Why the Sussexes are so angry in their entitlement. Where time makes money. And now, both our worlds: Rob Oseh on two women soccer players.
3. Too pleasant to be human – First, AI comes for your fast food. Would you make your dog go vegan? PETA wants you to. Plus lessons in knowing who you truly are. And not a tick-tock, but a Dick doc.
4. Trump gets thumped. On the economy…on Iran…by Heather Cox Richardson and Paul Krugman. Plus, The Old Man and the Strait.
5. Flips+Flops – How FIFA really makes money. What Gianni Infantino meant to say, but too late. Karl Marx and Gianni Infantino on who owns productive assets? And is artist anonymity a radical act in an age of self-promotion?
6. Health and welfare – First, Lyme Disease 101. And stayin’ alive. Can we have public health vending machines? And when can we get overnight inter-city trains? For those of us who believe power is the ultimate aphrodisiac, here’s a close second.
7. Found objects. Cassette tapes from around the world. Your own NBA team. Plus the downside of checking your phone 186 times a day, which you do now. Speaking of, where do iPhones cost the most? Plus an iceberg tips over in Greenland, and how to fix a cow. Finally, Caledon opens a 5-suite retreat.
8. Mark Gerretson and Danielle Martin. They’re both Liberal MPs and both use Instagram to devastating effect. Gerretson to attack Pierre Poilievre (on GDP and Dr. Theresa Tam) and Martin to praise the YMCA and Canadian health care. among many others.
9. Rules for writing. Stephen Pinker on how AI would rewrite his books. And Brett Stephens who says: “I’m begging you. Never write with AI.” Plus 7 Literary Sins.
10. What I’m watching: Lioness, Taylor Sheridan’s ultra-violent, ultra-fetching series on a fictional female U.S. Special Forces Team. Episode 1 of Season Three premiered this week, where we learn that the Ukraine War is like World War 1 meets Star Wars.
Also, I can’t remember the last time we binged a series, but we did this week on The Bombing of Pan Am 103, the 6-part Netflix dramatization of the bombing of a Pan Am flight from Heathrow to New York that was blown out of the sky over Lockerbie, Scotland, in 1988. Gripping and heart-breaking drama.
Guaranteed 100% written by Bob Ramsay, human.