New method aims to keep kids safe from illegal AI-generated content
AI Summary: Researchers, including Associate Professor Ashia Wilson and graduate student Vinith Suriyakumar from MIT, have developed a novel auditing technique to assess whether generative AI models can produce child sexual abuse material (CSAM) without generating illegal content. Collaborating with MIT's Healthy ML Lab and the nonprofit Thorn, the method analyzes the internal representations of models to identify adaptations that enable harmful outputs, achieving 100% accuracy in detecting CSAM-capable models. This approach provides a scalable solution for platforms hosting open-source models and law enforcement, addressing a significant gap in AI safety measures. The findings were presented at the "Trustworthy AI for Good" workshop at the International Conference on Machine Learning.