問題文
Which end-to-end evaluation confirms that a retrieval and generation solution works as intended?
選択肢
- Measuring both whether the correct passages are retrieved and whether the final answers are grounded and relevant, on the same set of questions
- Measuring the retrieval quality only, since it determines everything downstream and a generation step cannot invent an answer that the passages do not support
- Measuring latency and token consumption only, since those are objective numbers while any quality score depends on a judgment that can be disputed
- Measuring the final answer quality only, since that is what users see and a good answer means every layer beneath it must have worked correctly