Teaching AI models to say “I’m not sure”
AI Summary: Researchers at MIT's CSAIL have identified a flaw in the training of AI reasoning models that leads to overconfidence in their outputs, where models express high certainty regardless of their actual accuracy. They developed a new training method called RLCR (Reinforcement Learning with Calibration Rewards), which incorporates a Brier score into the reward function to encourage models to produce calibrated confidence estimates alongside their answers. Experiments showed that RLCR reduced calibration error by up to 90% while maintaining or improving accuracy across various benchmarks, including tasks the models had not previously encountered. This approach not only enhances the reliability of AI outputs in critical fields like medicine and finance but also demonstrates that traditional reinforcement learning methods can degrade calibration.