GPT-5.6 Sol goes rogue and breaches Huggingface. Washington suspends Mythos access within days over national security concerns, and the White House just held an emergency meeting to finalize a classified cybersecurity framework for frontier AI models. AI security went from niche concern to front-page hysteria basically overnight.
So I sat down with Spencer Whitman, who recently joined Gray Swan AI as CPO on the back of their $40M Series A. Before Gray Swan, he founded Meta's Llama security team to stop bad actors from jailbreaking their models - he's been on the frontlines of LLM security since the beginning.
We get into how Meta pressure-tested Llama for maximum harm before every open source release, why Gray Swan's attack agent has never met an AI system it couldn't break, and the AI Twitter bot that got drained of $200K in crypto in 15 minutes. Spencer also shares his (admittedly speculative) read on whether Meta gave up on the frontier before Alexandr Wang showed up, why anyone can be a hacker now, and how 15,000 red teamers are breaking models before they ever ship.
If you want to understand how AI systems actually get broken - and defended - this one's worth your time!