問題文
A support agent frequently answers a different question from the one the user asked. Which evaluator is the right starting point?
選択肢
- Task Completion, which measures whether a usable deliverable was produced
- Task Navigation Efficiency, which measures the efficiency of the tool call sequence and therefore also reveals when the wrong question was answered
- Tool Selection, which measures whether the right tools were chosen
- Intent Resolution, which measures whether the agent correctly identified what the user wanted