Quantifying Frontier Model Performance on Antibody Discovery Tasks
AI Summary: The Antibody Discovery Benchmark is a new experimentally grounded benchmark for testing AI agents' ability to make scientific decisions in therapeutic antibody discovery. The benchmark consists of 100 evaluations across ten areas of antibody discovery and was used to evaluate 20 model-harness configurations, with the strongest systems passing only about half of the attempts. The top-performing configuration, Anthropic's Opus 5 with the Claude Code harness, achieved a 53% pass rate, while GPT-5.6 Sol using the PI harness lagged behind with a 33.8% pass rate. Allocating more resources did not consistently improve performance, with some configurations achieving similar accuracy at lower costs and with fewer tool calls.