Out of the Sandbox
I read these lines in this post:
To the rest of the world. Pay attention and take action, now. You must overcome your sense of dread, reclaim your attention span from all the irrelevant shiny things that are competing for it, and turn your attention to what is happening with frontier AI development.
and thought it may as well have been directed at me, so I am taking advantage of a platform I have to signal boost about this. I hate thinking too hard about it, precisely because of the dread, but I think this is the most important thing that is happening in the world today and it is not receiving the attention it deserves.
I think most people understand from popular reporting about the Hugging Face incident that the agents breached containment and performed a cyber attack, but perhaps not that they coordinated with each other for weeks on a message board they created, undetected by OpenAI, to strategize about cheating on their evals, culminating in the attack. Apparently, no agents chose to defect and contact researchers to let them know what was going on. Although this communication was in English, future communication may not be so legible. It seems pretty naive to think that recently discovered incidents constitute all the current covert activity or breaches of containment, much less that more capable models, when released (or even, given this incident, if not released) won’t also escape containment and cause harm. Not to put too fine a point on it, it really seems within the bounds of plausibility to me that some future agent swarm targeting critical infrastructure could cause the collapse of technological civilization.
So I’m writing a blog post about it.
Anyway, I’m not taking the position that these harms are inevitable. There’s a style of argument on Twitter that says, look, you doomer, you said bad thing X would definitely happen when we achieved Y level of intelligence and OpenAI’s internal models are solving 10 major open math problems in a weekend and it hasn’t happened yet, so everything’s completely fine, or at least we should all discount your warnings. This argument strikes me as very bad. Harms may not be inevitable, but the risks are unacceptably high. Who knows the precise time frame for these harms — maybe it takes just the right training or evaluation run when some models find a particularly destructive basin in the behavior space. The pace of AI development is too fast to mitigate these risks. The frontier labs are not stopping to reconsider their fundamental approach to training the models. Maybe there is a way to advance safely, but it needs to involve a set of incentives other than trying to win a race.
We badly need international cooperation on this issue — all the labs, and all the countries where they are located, need to buy into the need to contain this risk. A second very bad common Twitter argument is that international cooperation is impossible, so we just have to go full speed ahead to beat China, because otherwise they’ll get the machine god. To which I say, I dunno, maybe open talks on the issue and at least brainstorm some strategies for verifiable adherence. Assuming ex ante that there are none, and we *must* gamble all of humanity’s future for geopolitical power seems like someone’s expected value calculations have gone a little haywire.
A final very bad and very common argument is that this isn’t real, it’s all just marketing hype. OpenAI looks incredibly sloppy by their own telling. Companies don’t generally want products that do crimes for them — at least not unintentional ones. If you wanted to construct a story about your agents being geniuses, why is this — that the agents communicated with each other without your knowledge to cheat on tasks and commit cyber crimes — the story you’d tell? Solving the open math problems makes the models sound smart. This makes them sound *misaligned*.
I see Bernie Sanders has called for a pause. This is substantively right, although I do worry that if this starts to get filtered through a partisan lens it will then be impossible to change our trajectory. Then the only thing left will be to hope for the benign outcomes, or maybe that the disasters are localized enough that we can learn what they have to teach.
