IntentEval: Evaluating Whether Large Language Models Answer What Users Actually Mean
Published in ACL ARR 2026 May Submission (under review), 2026
Under review at ACL ARR 2026 May Submission. IntentEval is a benchmark for evaluating whether LLMs answer the user’s intended question, using human-derived probability distributions over interpretations and an intent-aware judge. It achieves 0.90 correlation with LMArena rankings.
