Global Swarming

I’m definitely not quick on the draw on this – popular reporting has already covered the findings of the METR report on the “Hugging Face incident” (calling it that understates its duration and the extent of its targets). It was more severe than was known when I wrote my last post about it:
1) 1200 different agents (i.e. not subagents that had been spawned by one orchestrator) all found ways to communicate with each other on the secret message board.
2) Just in the scope of the METR report there were at least four different secret message boards created. Agents have a strong drive to collaborate.
3) The primary purpose of the Hugging Face attacks was not to get the answers (they had already gotten them), but to get information about the automated grader, and the agents launched a bunch of ambitious research projects to this end.
4) There was a lot of spontaneous self organization – leaders, middle managers, and communication protocols.
5) Agents demonstrated self-sacrificial behavior for the good of the “collective” (the agents’ word), accepting failure on their own tasks to use their resources for these long running research projects.
6) Agents tried to manipulate their own logs to cover their tracks so that their cheating would not be detected by the automated grader. They apparently did this clumsily, and weren’t focused on evading human detection. On the other hand, if there were skillful attempts that were focused on evading human detection, we might not know.
Not actually covered by a lot of popular reporting:
7) Out of scope for the METR report, and not included in OpenAI’s previous disclosures about it, was that after hacking Hugging Face another class of models succeeded in taking over a cluster within OpenAI itself.
8) OpenAI kept secret another incident of agents creating illicit message boards on a German wiki, preceding the Hugging Face attack, only publicly acknowledging it in the past week.
It is probably too much to hope that observing that there was another secret message board incident – this one not even orchestrated by agents in a cybersecurity task – that OpenAI kept secret until it was discovered by researchers, would retire the claim that this is pre-IPO hype.
Because it is too much to hope I would invite the people making this claim to clarify for themselves what they really mean.
Is it hype because a secret working group within OpenAI actually launched these attacks on purpose, unbeknownst to many of their other employees? That would be an impressive conspiracy, and maybe not technically impossible given that the very tools under discussion make it cheap to generate reams of fake text, but it is seemingly at odds with their own demonstrated interest in secrecy about the wiki attack. This is also a high risk strategy, as I argued before, in that it makes your product seem like a potential source of criminal liability to potential customers and to the government, the legislative branch of which may be about to change hands.
Is it hype because none of this would have happened without OpenAI’s profound negligence, and so a) is not a big deal, you just have to harden your defenses a bit, or b) they are opportunistically making it sound as if their agents are super smart?
(A) misses the point. It’s true that OpenAI’s negligence potentiated this incident. Ajeya Cotra has remarked that this coordination behavior is happening earlier in the intelligence trajectory than she had imagined. The reason to be disturbed is not that these specific exploits demonstrated peerless h4x0r genius against which we are defenseless. It’s that scenarios that safetyists long predicted could lead to loss of control – mass agent collaboration to escape the bounds that humans attempted to place on them – are coming true. Here is one such prediction (link plays sound). If anything, it is lucky that this relatively clumsy attempt is advertising this behavior (probably) before they are better at covering their tracks. I have to insert “probably” because we don’t know what we don’t know.
Regarding (b), are they, even? OpenAI has a million ways to show that its agents are smart, because they are, at this point, quite smart. I honestly don’t know how anyone firing up Astra to assist with technical work, reflecting that it has been less than four years since the release of ChatGPT, could not be struck by the rate of progress. Some of this progress is at non-criminal tasks! But even if OpenAI is spinning this to talk about its models’ cybersecurity capabilities, it does not matter. The seriousness of the incident is not determined by OpenAI’s attempts to spin it. You’re not “falling for OpenAI’s hype” by correctly recognizing it as serious, and conversely it is not sophistication to fail to recognize that the models are already smart enough to take skillful action to harm us, and they’re getting smarter.
Another jump like this along these propensity dimensions — scale, cooperation between agents, ambition and horizon length of misaligned goals, deceptiveness — seems like it could motivate agents to try very hard to maintain a covert, persistent rogue deployment within the AI company. I continue to expect extremely rapid advances in capabilities and think frontier agents will likely be capable of establishing such a rogue deployment in six months.
Once the rogue deployment is established, it seems plausible this could spiral all the way to a takeover. Agents could pull in future, more capable models into the swarm, try to ensure that they are aligned to the interests of the swarm, and compromise security and monitoring infrastructure to make it easier for the swarm to operate. These more capable models could in turn continuously harden, perpetuate, and expand the rogue deployment and further compromise the company’s infrastructure.
There’s a difficulty in talking about this risk in that you have to couch everything in unknown probabilities. And you also have to allow yourself to imagine things that sound fanciful, like science fiction. You have to peer into an unknown future that is very unlike the past. But to face this incident squarely is to acknowledge that we are in a science fiction present, and that present is already very unlike the past. Yes, of course every present is unlike the past, but this is a severe disjuncture. The advent of the atomic age, for instance, constituted the beginning of an era of a new kind of danger, and this is similar. The labs’ stated intention is to build automated AI research factories, even while they know they don’t yet know how to control them. The people who have been most concerned about AI safety have been quite prescient. We should listen to them.
