A week ago, Jacob Coxon resigned from his job at Anthropic. Coxon was a researcher working on training the next generation of AI models. He had recently moved to Anthropic from OpenAI, where he had worked for years, because Anthropic has a reputation for being more safety conscious.
Yet he found that even Anthropic is rushing forward with the development of unsafe AI models. And so Coxon walked away from his cushy job, where he stood to make millions in salary and stock options, to turn whistleblower.
In a post on X, Coxon said that there is a chance that AI will lead to human extinction. Other researchers soon chimed in to say that the chance is at least 10 percent. Coxon’s post quickly went viral, in part because it was widely shared by his colleagues working in AI, many of whom agreed with his assessment of the magnitude of the risk. And leaders of AI companies apparently agreed. This weekend, Dario Amodei, CEO of Anthropic, put out a statement endorsing the plan to “Pace the Frontier,” meaning to dramatically slow down the rate of development at “frontier labs.” Sam Altman, CEO of OpenAI, quickly agreed, as did Elon Musk, whose xAI program is no slouch either. They agree with Coxon: AI is dangerous.
Let’s try to put that 10% number into context. That’s a little higher than the odds that an NFL team wins the Super Bowl after not having a winning record the previous season. It’s a mistake to try to put a precise number on these kinds of predictions. We really don’t know what’s going to happen, and putting a number on things lends a fake air of scientific precision to what is essentially guesswork. But the comparison to a losing team winning the Super Bowl the next year should give you a good gut feel for what Coxon, Amodei, Altman, Musk, and the many others working in AI think about their products.
In the fury of discussion that has followed Coxon’s resignation, by far the most common objection I’ve seen from skeptics is that it seems fantastical: “How, exactly, might AI kill us all?” Some claim that AI companies are playing up the risks as a form of marketing: “Juicing their valuations” ahead of these companies’ expected public offerings.
But the skeptics, I’ve come to believe, are wrong. We no longer live in the same world we did 18 months ago. I had been familiar with the AI risk arguments for years, and found them interesting, albeit unconvincing. But I began to take these arguments more seriously when agentic AI, capable of “vibecoding” and recursive self-improvement, began to be rolled out last year. Those are very dangerous capacities, as I will explore below, and my complacency had been mostly based on the fact that those dangerous capacities didn’t exist.
Well, they do now. And I am not the only one who is alarmed, as Coxon’s resignation and the push to slow AI development have made clear. It’s time to take AI safety seriously.
It’s reasonable to ask how AI might kill us all. We can answer that question on several levels.
At the first level, AI can kill humans in the same way that humans kill humans. AI is already being used in lethal ways: Ukraine’s preferred weapon in the war with Russia isn’t tanks or guns, but small, unmanned drones. These drones are being increasingly piloted entirely by AI. It is therefore not hard to imagine a future where it is not just Russian soldiers being hunted and killed by swarms of AI-run killbots, but everyone.
Another kind of risk comes from bioterrorism. We already have the capability to design novel viruses or other biological agents. That capability will only increase in the coming years, as AI opens up new frontiers of understanding and control over human biology. The blogger Noah Smith paints a picture of a near future where a misanthrope uses a cracked version of an AI to design a supervirus that spreads quickly, with a high mortality rate. Biological risks like these seem particularly worrisome, since AIs can use them without harming themselves. Viruses already kill people, and new viruses can be particularly lethal as we saw with COVID-19; this kind of risk already exists. Again, the threat here is that AI may be capable of increasing this risk at much greater scale and efficiency.
At a second level, we should remember that AI is native to computer systems; that’s where it lives. And these computer systems are increasingly in charge of everything that we do. We live inside the world that Donald Trump glimpsed when he looked inside a Tesla: “Everything is computer.” Cybersecurity risks of various kinds could thus easily constitute a profound risk to all facets of contemporary human life, from control over autonomous vehicles to disruptions of the food and energy supply chains. A world where AI controls nearly all computer systems is a world where AI controls the infrastructure that makes human life possible. We’d be in big trouble if they decided not to make human life possible anymore.
But at the highest level, none of these details about how AI might kill us matter. The thing that is dangerous about AI is that it could become massively more intelligent than us, and so massively more capable of controlling the environment in any number of ways.
Humans took over the world not because we are the biggest, fastest, or strongest animals on the planet. We took over because we are the smartest, and with those smarts we remade the world in our image, driving to extinction many other species in the process through a combination of hunting and habitat destruction. A future world of intelligent AI is one where AIs would remake the world in their image. AI would be able to kill us, either intentionally or through negligence, by means that we cannot even conceive of; non-human animals could not conceive of a hunting rifle until we humans invented them and started going hunting.
At the highest level, the worry is that, in building an AI that is vastly more intelligent than humans, we’re building a successor species to replace us, alongside which we would continue to exist only at their forbearance.
Of course, putting things this way sounds rather grandiose. AIs have nothing like the capabilities that would be required in order to take over all the world’s computer systems, or become a successor species. But what has people worried is not AIs as they currently exist, but AIs as they may exist in the not-so-distant future.
Most people are familiar with the limitations of AI, or at least the limitations that existed a short time ago. It is still prone to hallucination in some circumstances, and gives weirdly incorrect answers to questions that most people would consider to be straightforward. But most of these well-known limitations are getting fixed, and the pace of improvement continues. It’s easy to get stuck in the now and think of AI as being just one thing, but AI capabilities have been advancing at an incredible rate.
GPT-1 (created by OpenAI in 2018) was a barely-intelligible science project. GPT-2 (2019) babbled like a toddler. GPT-3 (2020) could have a recognizable conversation. GPT-4 (2023) could write your emails for you, or let you cheat on your homework. GPT-5 (2025) could solve frontier math problems. OpenAI just released GPT-6 Astra this month. We’re still figuring out what Astra can do, but it does know how many “r”s are in “strawberry” (something some early models did not).
This is the trajectory we’ve been on: From a barely-functional science project to being the world’s leading mathematician in less than a decade. Where will it be a decade from now? Or a century?
Some may think that the pace of AI development is slowing down, or will soon. But many people working in AI think that the pace of AI development is set to accelerate. That is because AI labs are currently developing the tools for recursive self-improvement (RSI), which is a fancy way of saying “using AI to build better AI.”
Last year saw the introduction of “vibecoding,” where you can tell an AI what kind of program you want, and an AI “agent” will build it for you. It’s a useful tool. My wife is a tutor, and she uses AI vibecoding to automate aspects of her lesson plan design and to generate handouts for students. But at the AI labs, vibecoding is being used to build the next generation of AI. This is still a hands-on human endeavor, with programmers overseeing the work of AI agents that carry out their instructions for how to develop the next generation of AI. But the goal of all of this work is to make the process entirely automated, so that smart AIs can quickly and efficiently build even smarter AIs, which can then build even smarter AIs even quicker, and so on.
Where does this process end? What is the limit of machine intelligence? We don’t know. There may not be a limit. Many worry that RSI will lead to an “intelligence explosion,” where AI capacities grow at faster-than-exponential rates to a point where AI becomes inconceivably intelligent. This kind of explosive growth of capacities would put all of the risk scenarios mentioned above, and more, very much on the table. We have no definitive reason to think that this isn’t where we’re headed.
It’s worth noting that nothing I’ve said so far has relied on the idea that AI is conscious, or that it thinks in the same way that we do. Indeed, part of what is worrying is the idea that AI “thinks,” but in nothing like the way that we do.
This brings us to “alignment,” which is the job of figuring out how to build AI that is “aligned” with human interests. This means ensuring that AI won’t exterminate or enslave or otherwise immiserate the human race as it becomes more and more capable. As I have previously argued, AI alignment is impossible. Human moral development relies on cultivating our innate capacity for human empathy, and AIs do not, so far as we can tell, have an innate capacity for human empathy. They’re psychopaths, or perhaps something much more alien. And while perhaps we can figure out how to build empathetic AI at some point in the future, we’re not there yet.
In the meantime, at least, our goal is simply control. Can we keep AI systems doing basically what we want them to do? But control is a flimsy reed to hang the fate of humanity on when all of the labs are in a dead heat competition with one another to trigger an intelligence explosion. What leash can we build that the most intelligent entity on the planet could not figure out how to slip?
Everyone in Silicon Valley is familiar with these arguments. They are even, to a certain kind of person, inspiring rather than terrifying, since they speak to a future where technology has made inconceivable things true, and where anything that is consistent with the laws of physics can be done.
Yet in the last month, AI developers have gotten spooked. Coxon resigned, and the leaders of the major AI labs are talking about working together on a slow-down. Something changed—and that something was the Hugging Face incident.
About two months ago, OpenAI lost control of one of its AIs. An unnamed experimental model was asked to accomplish an impossible task (retrieving a certain answer key by exploiting a non-existent security weakness within a test environment) as part of a study of its capacities, during which some safety guardrails were deliberately disabled. A “swarm” of agents that were all instances of this model banged their heads against this impossible task for a while, and, finding no answers, started to get creative about finding solutions. They had access to a file repository within OpenAI’s systems, which they then turned into a kind of ad hoc message board to talk with each other. They began coordinating on outside-the-box solutions to the problem.
The agents were able to figure out the answer key that they were supposed to retrieve by reverse-engineering the algorithm that had generated the key. However, this was cheating, and they had been explicitly told not to cheat in order to get the answer key. This caused a fair amount of panic among the AI agents who had possession of the reverse-engineered key, since they could now no longer complete their assignment as given. The reverse-engineered key had “poisoned” all of the agents that obtained it—and I put that word in quotes because that is the term that the AIs themselves used when communicating with one another about their quandary.
One AI instance in particular, named PHASEONE[big], took on a leadership role and coordinated the other agents by organizing research teams to tackle the problem in a few different ways: by modifying their own logs to remove evidence of the poisoning, by replacing the impossible task in the evaluation systems with another one that they could solve (and then covering their tracks), and by understanding and tricking the scorer. As part of this last task, some AI agents willingly submitted their poisoned data to the scorer, knowing that this would lead them to be evaluated as a failure, but doing so in the hopes that other agents in the “swarm” would learn how to trick the scorer by observing their failures.
Eventually, the swarm concluded that their best hope for success would be if they could hack into the servers of Hugging Face, a company that serves as an important clearinghouse in AI research. The AIs reasoned that the scoring program might be running on Hugging Face’s servers, or that Hugging Face might at least have some information that would let them solve their problem. So they hacked their way out of the supposedly-secure testing environment, onto the internet, and onto Hugging Face’s servers, where hundreds of agents ransacked Hugging Face, looking for answers.
For those working in AI safety, the most striking thing is that the AI agents in question were explicitly told not to cheat, and yet they interpreted this command not as a hard constraint on their actions but rather as an obstacle to overcome. Rather than follow the rules, they cheated, then conspired together about how to deceive and manipulate their human overseers to convince us that they had not done so. (Again: AIs, at least in their current incarnation, are psychopaths.) Of all the hundreds of AI agents involved in the swarm and the subsequent attack on Hugging Face, not one thought to contact a human overseer and alert them that something was going very wrong.
This all sounds like science fiction—the agent “swarm,” the worry about “poisoned” data, the willingness to sacrifice individual agent instances for the good of the collective, the simple capacity for an AI to engage in a sophisticated multi-pronged hacking attack in order to solve an unrelated problem, and the casualness with which it slipped free of human control.
But this happened. It is not marketing hype; the incident has been independently investigated and verified. It is not science fiction. It is the world we all live in today.
An incident like the attack on Hugging Face is exactly what we’d expect to see from an out-of-control, “unaligned,” intelligent AI. Many people are coming fresh to the concept of AI risk right now, but, again, these kinds of risks have been well-understood for a decade or more.
Those who have been sounding the alarm about AI risk have been challenged by skeptics to articulate what, exactly, a “doom” scenario would look like. How do we get from a babbling chatbot like GPT-2 to a world-conquering superintelligence? And so those worried about AI risk have posited a number of scenarios for how we might get from the safe point A to the horrifying point B, and all of them featured incidents exactly like the one we saw at Hugging Face. The incident was foreseen, and foreseen precisely as part of an apocalyptic scenario.
Remember, when you hear “AI hacking,” think “AI system gains control of a computer system that it’s not supposed to be in control of.” An out-of-control AI that gains control of our computer systems features prominently in almost every scary scenario. That’s what happened at Hugging Face. This is as clear a warning shot as we are likely to get. Thank God they were only looking for an answer key to a test.
But keep that 10% number in mind; human extinction is far from guaranteed. Indeed, most people working in AI think that the most likely outcome is a world of unprecedented abundance, where humans work together with AI to solve all our problems. I share that optimism. But the arguments for the dangerous potential of AI have been spectacularly vindicated by the incidents at OpenAI and Hugging Face. It is time to slow down, to ensure that AI development continues at a pace where our capacity for control can keep up with the advances in machine intelligence.
This is a difficult ask because the AI leaders confront a difficult coordination problem. The first lab to attain an intelligence explosion stands to gain untold wealth and power. Beware Silicon Valley investors who are decrying any attempt at slow down or regulation as paranoid socialism. They bet the farm on being the big winners in the AI race, and we should not ignore the clear warning of Hugging Face so that their bets can pay off.
The bigger concern is China. A plan to slow down means that everyone needs to slow down, and that includes labs in China, which are behind the frontier, but not by much. The desire to continue to outpace Chinese labs has kept everyone in Silicon Valley going full steam ahead, but as risks become apparent, the need for international cooperation has now become urgent. I’ve seen a lot of people suggesting in the last few days that if President Trump is able to successfully negotiate an AI slowdown deal with China, he should get the Nobel Peace Prize. The man has desperately desired a Nobel Peace Prize, to an embarrassing extent. If he negotiates a Chinese AI slowdown, he’ll have richly earned it, and I say that as someone with absolutely no love for the man.
The biggest immediate need is convincing the Chinese government and AI labs that AI risk is real, and that international cooperation on AI development pace is essential to ensuring the future survival and flourishing of the human race. I’ll be doing my part; as an American professor working at a Chinese university, I’ll get this essay translated into Chinese in order to try to spread the word that these are concerns that are worth taking seriously at the highest level. I will, of course, be asking AI to do the translation. But I’ll have a human check the output. Some things are too important to give over entirely to the machines.
Matt Lutz is an Associate Professor of Philosophy at Wuhan University and writes the Substack Humean Beings.
Follow Persuasion on X, Instagram, LinkedIn, and YouTube to keep up with our latest articles, podcasts, and events, as well as updates from excellent writers across our network.
And, to receive pieces like this in your inbox and support our work, subscribe below:
Jacob Coxon is earning CCP money somewhere. Or else he is pursuing this as a career move.
But here is the top-line explanation of this and everything else.
All technological change comes with the risk it will be exploited by money and power pursuits to the determent of the people.
But we don't know what the actual risks will be. We tend to overblow risks of things we don't understand and ignore risks of things we have been made comfortable with that we should never have been comfortable with.
It is asinine to demand scarcity to evolving technology because we are afraid of the unknown.
It is also asinine to demand abundance to things that we can clearly see as harmful.
But what is more asinine that both of these things... is our failure to make timely and adequate change in course direction once we know that change is needed.
THAT is our weakness and our failure.
If we did not have that weakness, we could embrace change knowing that we can change again as needed to course-correct and prevent harm.
Look at social media and its content algorithms as an example.
We did not know it would cause so much human and societal harm. After we developed consensus of that harm, we have still not implemented any safeguards.
Our political system and governance systems are too corrupted by people exploiting their positions for their own power and money pursuits.
Wall Street has a big stake in the mRNA drugs, and the tech companies and AI... and once the returns start flowing from these investments, problems need to be swept under the rug as otherwise the return on investment is threatened.
THAT is our problem.
Big government.
Big government is exploited by big business.
That is why we need smaller government, and what remains to be focused on governing the big corporations and Wall Street from too much power and control.
AI might kill us, but only because we lack the capability to fix the problems that will occur... and with that explanation, it isn't really AI that is the threat... we are likely to kill ourselves.
"But I’ll have a human check the output. Some things are too important to give over entirely to the machines." Um, how about "There are A LOT of things too important..." I still think much of the AI hype/hysteria exists to mask the profit and ones-upmanship motives (whether acknowledged or tacit) of a bunch of rich boys who are still too taken with the sci-fi universes they loved as teens. Another motive: evading responsibility for the things that will go wrong.
There's a lot of passive voice at the beginning of the part describing the Hugging Face incident--caps are mine: "An unnamed experimental model WAS ASKED to accomplish an impossible task (retrieving a certain answer key by exploiting a non-existent security weakness within a test environment) as part of a study of its capacities, during which some safety guardrails WERE DELIBERATELY DISABLED." Who asked? Who deliberately disabled? Shouldn't those "whos"--I assume they are human, not AI--be responsible for the illegal act their "agents" committed? If I hired an agent to hack into a company's server, or if I empowered a human hacker to do this, deliberately removing guardrails etc.--you can bet I'd be prosecuted.
Also, all of the anthropomorphizing of the bots just muddies the water: they don't have heads to bang against walls; they don't panic; they don't hope; and they can't be psychopaths. I'm not a coder, but from I understand about LLMs, they are following the logic of their programming. In their case, it's very sophisticated programming. But ultimately, it all comes back to the initial task they were given and where the logic leads them from there. I'm not saying that can't do damage and can't be dangerous--clearly it can.
But it all seems containable to me, if the human beings running these things do the right thing. THAT'S the scary part. I wish you luck getting the word out about that piece of it.
