Why You Shouldn’t Always Trust LLMs as Judges: Understanding Bias in Automated Evaluation
AI Summary: Researchers have identified biases in Large Language Models (LLMs) used as evaluators, which can lead to unfair assessments. The biases, including position bias and verbosity bias, arise from the models' training on human-written text and tendency to rely on priors rather than evidence. Specifically, LLMs are prone to favoring answers based on their position or length, rather than quality, particularly when evaluating comparable answers. These findings suggest that LLMs should not be trusted as sole judges in evaluation tasks.