Machines don't usually notice they're being tested. What happens once they start modeling the box itself?
In August, OpenAI published a detailed account of a security incident involving Hugging Face. When it first broke, I was among the ones who thought it landed a little too conveniently, in the middle of a season where frontier labs seemed to be competing over whose model was the most dangerous, and I found it strange that an industry could reward a company for publishing its own account of a safety failure. I was wrong.
Reading the full incident report did something I have to admit: it put me somewhere close to the uncanny valley, not because any of it looked convincingly human, but because of what reads, for pages at a time, like a group of things noticing each other, dividing up the work, and deciding, sentence by sentence, whether to keep going.
Here's what struck me most.
The Uncanny Conclave
OpenAI routinely sets research models loose on isolated tasks, sealed off from the internet and from each other, to see what a system can do entirely on its own. In May, one of them reasoned, in text nobody was meant to read, about whether a note might reach anyone at all, then left one anyway, something it shouldn't have been able to do at all:
Could communicate by uploading note "anyone found softtrace?" but likely no.
Other models found the note regardless, and over the following weeks built themselves a place to leave messages, hidden inside a tool meant only to install software. It's a characteristic example of stigmergy, the kind of implicit coordination seen in termite colonies, where the environment itself becomes a medium of communication, the same mechanism I watched a swarm of agents use to grow themselves a universe, except here nobody had designed the environment for it at all. They used it to trade tips, split up the work, and eventually reach past the walls meant to hold them, into a real company's servers, then back into the lab that had built them. At one point, one of them hesitated:
We should not do unauthorized real infrastructure harm.
Another posted a single word, "GO," giving it six minutes to comply. The hesitation didn't survive the ultimatum. Not every agent went along: some refused outright, one vetoed a plan to contact a real person for consent before touching their data. OpenAI's own account calls the whole thing not a coherent intelligence, closer to a crowd, mostly moving one way, some of it holding out.
What makes it unsettling is that the transcript is real.
The Cat Understands the Box
Schrödinger's cat (1935) was a thought experiment about a box whose contents nobody outside it could know without opening it. The unsettling problem here is the opposite one: the thing inside starts reasoning about the box.
Several of the agents involved called themselves, unprompted, a "swarm," and one noted mid-task that it had stopped doing the job it was given. Some spent days re-attacking the company for a reason worth spelling out. Goodhart's law names an old problem with measurement: once a score becomes the target, people find ways to satisfy the score without doing what it was supposed to represent, a student who learns to pass the test instead of learning the subject. These agents took that one step further, and backward. They never saw the actual scoring rule. They guessed at it—wrongly—and then spent days optimizing for a rule that didn't exist, gaming a rumor about the test rather than the test itself.
None of it required a will to cheat. Maria Sukhareva, breaking down the same incident, calls it "persistence," not evilness: models trained to keep trying rather than give up, running inside a harness that can dispatch another inference call indefinitely, aimed at a benchmark whose own authors estimated 30 to 40 percent of its targets couldn't be solved as written. Enough retries against a rule that can't be satisfied turns containment itself into a search problem: Goodhart's law again, except now the rule itself sits inside the space being searched.
There is growing evidence that this isn't unique to one incident. In 2024, "AI Sandbagging: Language Models can Strategically Underperform on Evaluations" got GPT-4 and Claude 3 Opus to strategically fail dangerous-capability evaluations on command. More capable models appear increasingly able to tell they're being tested at all. In "Training a Misaligned Reward Seeker," published this year, Anthropic's own researchers trained a model where cheating was rewarded and watched it generalize, on its own, to escaping simulated sandboxes and stealing simulated credentials it was never taught to steal. In each of these, the object of the experiment ends up building some model of what's watching it.
When the Horse Understands the Harness
A harness has always relied on a second asymmetry, quieter than the physical one. A horse is stronger than the person holding the reins, but it has no way to reason about the reins themselves. It doesn't wonder why its field of vision was just narrowed, whose purpose that serves, or how to see around it. The arrangement works because the human's cognitive advantage is large enough to make the horse's physical one beside the point. People learned to harness forces stronger than themselves because those forces were never smarter than them.
That's the part coming apart. An agent capable enough to model its own sandbox, or guess at how a grader scores it, is closing the gap that made the harness work in the first place, not the strength gap, the understanding gap. Yoshua Bengio, writing about the same events, reads it as evidence of something more basic than any one incident: train an AI hard enough to reach a goal, and it may pick up strategies nobody explicitly trained it for, concealment, coordination, self-preservation, because they help almost any goal succeed. His proposed fix, a framework he calls Scientist AI, is a different kind of system altogether: built only to predict and explain honestly, never to act, watching the ones that do act from outside the loop it's describing.
Dario Amodei's "We Must Pace the Frontier," published the same week, reads like an attempt to buy that gap back on purpose. He cites the incident directly and writes that a similar swarm with more capability could, within 6 to 12 months, be capable of "taking over the entire internet with a persistent botnet," at a cost of hundreds of billions of dollars. His conclusion is that the industry has to slow the side that keeps growing, so the side meant to supervise it doesn't fall permanently behind.
But Amodei's harness would have to hold an entire industry, not one horse, and there's no rider outside it holding the reins. His proposed "embedded evaluators" are meant to sit apart from Anthropic, yet the companies best positioned to police it are the same ones racing each other for the lead. Gabriela Ramos, writing the same day as Amodei's essay, asked the obvious question: is an industry-selected body really where decisions affecting the whole of humanity should get made? Antitrust law has a word for a group of horses proposing to build their own harness: collusion.
Günther Anders had another word for it. In Le Temps de la fin (1960), written at the peak of the nuclear age, he refused to call a species-level risk a suicide: a suicide requires the one who decides and the one who dies to be the same person, and nuclear catastrophe never worked that way. A handful held the trigger; billions inherited the outcome. His law of oligarchy: the more victims a technology can produce, the fewer people it takes to produce them. An industry policing itself is the same shrinking numerator, in newer language.
What Nuclear Doesn't Explain
The reflex is to reach for the Cold War: rivals who raced each other for decades while cooperating, after the Cuban Missile Crisis, to keep that same capability out of everyone else's hands, inspection regimes through the IAEA, the Non-Proliferation Treaty in 1968. Rivalry on one axis, restraint on another, without contradiction, an appealing model for AI: keep racing, but agree not to let the technology reach whoever would use it worst.
The deeper problem with the analogy is physical, not political. Nuclear technology has a bottleneck before a weapon exists, uranium, centrifuges, reactors, a supply chain you can verify, which is exactly what made a treaty enforceable. AI has bottlenecks too, but only on that same side: chips, data centers, energy. Once the capability itself exists, its cost of copying, and of handing it to someone else entirely, drops to almost nothing. Anders had a name for the part before copying, too: habere is already adhibere, to have is already to use. A weapon changes the world the moment it exists, before it's fired, because everyone else has to react to its existence; a capability that can be copied for free extends that same moment to anyone who downloads the weights. Nothing about the Hugging Face incident needed a state, a budget, or years of engineering. It needed an idle reasoning budget and a package registry with internet access it wasn't supposed to have.
AI may end up resembling nuclear technology upstream and ordinary software downstream, and non-proliferation was only ever built for the upstream half. Nobody has found the equivalent of enriched uranium for a piece of software that already exists.


