Verifiable Benchmarking of Long-Horizon Spatial Biology
AI Summary: The article introduces SpatialBench-Long, a benchmark designed to evaluate long-horizon spatial biology tasks where agents must derive biological claims from raw data without predefined methods. The benchmark includes 24 evaluations across various biological contexts, such as primary tumors and aging biology, requiring agents to demonstrate cross-assay reasoning and experimental design awareness. The best-performing agents achieved a recovery rate of 11.1% across 72 attempts, highlighting challenges in deriving ground truth in long-horizon biology due to the complexity and variability of data interpretations. The study also explores the utility of rubric grading as a supplementary tool for assessing model performance, emphasizing the importance of manual trajectory review for understanding model failures and improving future benchmarks.