Abstract
As AI systems grow more capable, the potential for extreme misuse by non-state actors, particularly involving chemical, biological, radiological, and nuclear (CBRN) threats and cyberattacks, has become an urgent concern for governments, companies, academics, and the public.
This report maps five current safeguard practices across three levels: two model-level interventions (pre-training data filtering, and post-training techniques), two deployment-level interventions (monitoring and access control), and one governance-level intervention (pre-deployment evaluations) that informs deployment decisions (including release conditions and safeguard design) and supports their ongoing refinement thereafter. For each intervention, we describe current practice, implementation across frontier developers, empirical evidence, limitations, and priorities for future research and collaboration.
Authors
Isabella Duan, Edward Kembery, Stephen Casper, Zhikai Chen, Jasper Götting, Geng Hong, Zaheed Kara, Lijun Li, Xiaojian Li, Kellin Pelrine, Ziyue Wang, Jia Xu, Ziwei Xu, Zhiyuan Zeng, James Zhang, Jie Zhang