問題文
A government agency is building an evaluation set for its records assistant. Most questions in production are routine lookups, but the costly failures happen on rare question types: multi-record comparisons and questions about superseded regulations. Which two principles should shape the evaluation set? (Select two.)
選択肢
- Match the evaluation set's composition exactly to production frequency, so the aggregate score predicts production performance.
- Exclude the rare types from evaluation and handle them with human review instead.
- Report results per question type rather than only as a single aggregate, so that the over-representation does not distort the overall figure's interpretation.
- Keep the evaluation set small so it can be re-run frequently during development.
- Deliberately over-represent the rare, high-consequence question types relative to their production frequency, so that enough cases exist to measure performance on them.