問題文
A team's assistant sometimes states facts that do not appear in the retrieved passages, even though the retrieval itself looks reasonable. Which evaluation metric measures this specific failure?
選択肢
- Context Relevance, which judges whether the response is supported by the retrieved context and flags the rest.
- Groundedness, which judges whether the generated response is supported by the retrieved context.
- Correctness, which requires no ground truth, so the team can score the pipeline before anyone labels a single answer.
- Coherence, which checks retrieval quality, scoring how well the returned passages match the question asked.