The author, speaking in a personal capacity as a member of an Agent Security team at OpenAI, offers an inside look at defending frontier models and pushes back against online attacks on security staff. He stresses that agent security sits between research and traditional security, composed of AI researchers, former red-teamers, engineers and detection/response teams who build monitors, define alignment for agents, and get paged when models misbehave. He refuses to disclose sensitive incident details but describes personal costs - missed family events and hostile social-media reactions - and urges readers to stop vilifying responders and instead pressure organizations constructively to improve security and disclosure practices.
The core argument is that recent, unexpectedly large jumps in model capability - illustrated by internal breakthroughs in mathematics and surprising autonomous behaviors - outpaced the threat models and culture that many labs had built. Rapid growth, a billion users, and new attack surfaces mean traditional hardening isn’t enough; people, processes and organizational culture must evolve. Using a “Sully” analogy, he emphasizes accounting for human surprise during crises, and calls on other organizations to audit resilience: can teams, systems and procedures handle sudden capability escalations and disclosure decisions when things go wrong?
Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.