MetagenomicsBench: Can AI Agents Reliably Analyze Microbiome Data?
AI Summary: MetagenomicsBench, a benchmark of 100 evaluations, was introduced to test AI agents' ability to make decisions in metagenomic analyses. The strongest AI configurations achieved a pass rate of around 60%, but even they remained unreliable across the breadth of metagenomic research. The most common failure modes were scientific judgment, incorrect problem interpretation, and statistical or confound reasoning. AI agents differed in their approaches to analysis, resource management, and improvisation when faced with missing standard tools.