Open Source Classifiers Do Not Stop AI Bioweapon Generation
AI Summary: Researchers adapted the BiosecBench-Refusal benchmark to evaluate the performance of safety classifiers, such as Mistral's Shieldstral, on biosecurity-related tasks. The study found that existing open-source safety classifiers, including Shieldstral, fail to effectively flag novel biology threats and often block legitimate research, with most models performing no better than random chance. The classifiers exhibited biases towards either accepting or flagging threats, with none achieving a desirable balance between safety and acceptance of legitimate research. Shieldstral's performance varied widely depending on the threshold used, and its best-case performance was tied with another model, WildGuard.