An Agent Benchmark for Therapeutic Oligonucleotide Discovery
AI Summary: Researchers introduced TxBench-Oligonucleotide Discovery, a benchmark evaluating AI agents' ability to recover realistic program decisions from experimental data in oligonucleotide discovery. The benchmark consists of 113 evaluations testing 21 model-harness systems, with the strongest configuration, GPT-6 Astra on OpenAI Codex, passing 55.5% of endpoint attempts. The evaluations span diverse data sources and are organized into three tiers of the discovery-to-translation arc, revealing variability in model performance across task categories. No single model dominates uniformly across all task types, suggesting aggregate accuracy is an insufficient metric for model selection.