Why it’s critical to move beyond overly aggregated machine-learning metrics
AI Summary: MIT researchers have highlighted significant failures in machine-learning models when applied to new data environments, emphasizing the necessity for rigorous testing before deployment. Their study, presented at NeurIPS 2025, found that models trained on data from one hospital could perform poorly on up to 75% of patients at another hospital, despite high average performance metrics. The researchers identified that spurious correlations, such as irrelevant markings on X-rays, can lead to unreliable predictions, particularly in medical diagnostics. They introduced an algorithm, OODSelect, to detect instances where the expected performance order of models does not hold across different settings, challenging the assumption of accuracy-on-the-line in model evaluation.