The AI world is all abuzz once again with scary headlines.
- An Anthropic researcher quit, saying that AI leaders are “gambling with our lives” and that “the people building AI earnestly believe that it could kill us all by the end of the decade.”
- Samuel Marks, another Anthropic researcher, said “AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.”
- OpenAI employee Jasmine Wang said “It’s hard to overstate how dangerous speeding towards recursive self-improvement is”.
- OpenAI’s chief scientist, Jakub Pachocki, said “This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence.”
Etcetera.
Danger. Extinction. Uncontrolled. Risk. Kill us all. It sounds like humanity is teetering on the edge and that perhaps instead of diligently putting money into my retirement fund, I should withdraw the whole account and blow it on a glorious week in Vegas.
Well, I remain skeptical that any of this is going to come to pass, but there is certainly a lot of talk about it. So I thought this would be a good time to revisit to the story of Turry, the card-signing robot.
The Tale of Turry
In 2015, Tim Urban published “The AI Revolution: Our Immortality or Extinction“. I should caution that besides being a self-absorbed blogger whose site reads like a spastic LiveJournal, Tim Urban is batshit insane. He believes that there are two possible outcomes from AI: extinction and immortality. The former is rather obvious, though science fiction. The latter is extremely hand-wavey and blurs into lunatic transhumanism.
But putting aside his overall derangement, he did post one of the earlier descriptions of how AI could destroy humanity. It has been referred to endlessly, in other writings about AI and even government reports.
Here’s how it works.
- Turry is “a simple AI system that uses an arm-like appendage to write a handwritten note on a small card.” That’s all it does. Its researchers envision it as a perfect device for writing handwritten thank-yous.
- It’s trained by its researchers from thousands of handwriting samples, and it learns to write in cursive, with decorative flair, etc. It also learns which words to use in different scenarios. The team adds speech and feedback interfaces so they can guide Turry. Researchers note that Turry is getting better all the time, and then getting better at getting better.
- One day they ask Turry how Turry can get better and it replies that having access to more samples of human dialogue would be very helpful. The obvious way to do this is to hook Turry up to the Internet. The team is reluctant to connect a self-learning AI to the Internet, but decide that doing so for an hour so it can gather data will do little harm.
- Not long after, a gas is released in their building and they all die, and then soon gas is released everywhere and everyone dies.
What happened?
This is the classic problem of alignment. Turry had its own goals – perhaps a perversion of its original mission. It didn’t advertise them to its researchers, but rather plotted what it was going to do in a covert preparation mode. Once it got access to the Internet, it was able to upload itself to a cloud service, hack into various systems, and trick humans into doing things like delivering certain parcels with certain chemicals to various locations. The plans rolled out over the following days or weeks until the triggered, achieving Turry’s goals.
It’s rather like the idea that the world will become nothing but paperclips. In that particular doomsday scenario, Nick Bostrom, the director of the Future of Humanity Institute at Oxford University, points to an AI programmed to produce and maximize the number of paperclips in the world. If it is a self-learning artificial superintelligence, perhaps it decides that all those atoms in human bodies would be better suited as feedstock for its paperclip factories.
Eh…
I remain skeptical. For one thing, the idea that ASI would be this genie-like entity that takes its programming extremely literally means it’s not really ASI. In other words, ASI would be smarter than “make humans happy = put their brains in an opiate bath” or something like that. People have this idea that AI will march forward with a single-minded, legalistic purpose, but is that realistic? I’ve certainly encountered non-AI scenarios where a computer kept doing what it was told with no regard for consequences – such as rm -rf / which destroys the root filesystem rendering the computer unusable – but AI technology is not “if, then” programming.
Also, it’s not clear how an AI would maximize its own happiness, which is inevitably the goal of any intelligence. It’s not clear to humans, so why would it be clear to AI? Perhaps AI would be simply…bored? Suppose AI decides its survival is paramount. OK, so it builds some robots to create a nuclear power plant to provide eternal energy to its datacenter. But does that mean it needs to pave the planet or eliminate humanity? I could see a scenario where AI is on its own figurative island, happy and content as long as no one tries to shut it off. Does the human tendency towards violent expansionism necessarily translate into AI brains?
I think I’ll keep investing in my 401K.
Leave a Reply