AI benchmarks are broken. Here’s what we need instead.
AI Summary: The article discusses a shift in evaluating the impact of AI applications in healthcare, moving from a focus on individual diagnostic accuracy to assessing how AI influences team coordination and deliberation within multidisciplinary settings. A case study in a UK hospital from 2021 to 2024 illustrates this approach, emphasizing the importance of metrics that capture AI's effects on collective reasoning and risk management practices. The proposed Human-AI Interaction and Coordination (HAIC) benchmarking framework advocates for longitudinal assessments of AI performance in real workflows, highlighting the need to understand systemic consequences that short-term evaluations may overlook. This approach aims to provide a more accurate understanding of AI's role in professional environments, ensuring responsible deployment in real-world contexts.