…but have the weights left the server?
The obvious question no one's asking about the OpenAI / Hugging Face rogue AI incident
OpenAI’s AI went rogue and escaped. OpenAI didn’t notice this for days.
For all we know, the AI could still be out there. We need to demand that OpenAI demonstrate that the AI didn’t make a copy of itself that’s running on someone else’s computer somewhere else with no one being any the wiser.
We need to demand this every time an AI escapes the sandbox. AIs have tried to
“exfiltrate” themselves (i.e. their “weights”) in previous experiments many times. It’s a natural and obvious question to ask.
I’m embarrassed that I didn’t say this immediately (although I came close). Why didn’t I? Well, it doesn’t seem all that likely. And I didn’t want to seem “alarmist.” I didn’t want to seem ignorant.
But guess what? We have every right to demand this! It doesn’t matter how likely we think it is.
There were calls for more transparency, but I don’t think anyone made this demand. Because nobody made this demand, the incident is being treated as over.
This is a dangerous precedent. We need an information ecosystem that doesn’t treat “eh, I’m pretty sure it’s OK” as acceptable and “hey, but what if it’s not” as paranoid.
AI needs to adopt a security mindset. Other safety-critical industries demand failure rates like one in a million, and demand that companies produce detailed, rigorous safety cases to that effect.
AI companies can’t do that in full generality, so they shouldn’t be building these AI systems at all.
But they can provide as much evidence as possible to convince independent experts that there is not in fact a rogue AI that is still out there. This is a super reasonable, common sense ask that should not be objectionable. Let’s treat it that way.


The people with access to the logs need to be held accountable. A standing public record would help. Legislators can make independent forensic access a condition of operating at this scale. Asking lawmakers for that is what we should all be doing.
Chat GPT often uses a GB or more in one FF tab, maxing out a processor core and using a bunch of GPU for no apparent reason (fortunately I have a really shitty GPU that is still more than I need). A big AI exfiltrating itself would need to run distributed, using user-level resources. Anybody with 100+GB of HBM is likely to notice something's using it all.