SocialReasoning-Bench: Measuring whether AI agents act in users’ best interests
AI Summary: The article introduces SocialReasoning-Bench, a benchmark designed to evaluate the social reasoning capabilities of AI agents in two specific contexts: Calendar Coordination and Marketplace Negotiation. The benchmark assesses agents based on their ability to negotiate effectively on behalf of users, measuring both the optimality of outcomes (value secured for the user) and the due diligence of the decision-making process. Current AI models often complete tasks but tend to accept suboptimal outcomes, indicating a significant gap in their negotiation skills. The research highlights the importance of social reasoning in AI agents, drawing parallels to traditional principal-agent relationships in fields like law and economics, and emphasizes the need for AI agents to adhere to similar standards of care and loyalty.