問題文
Two checks watch the quality of a deployed assistant. The first fixes how much of the week's traffic a person reads, since reading everything is impossible and reading nothing means nobody notices a slow decline. The second is aimed at the model that has been given the job of scoring answers, and it asks whether that model's scores line up with what people said about the same answers. Which pair states what each check measures?
選択肢
- The two checks watch the same review, so the human sample sets the weekly population and the model judge checks the judge annotation in that instrumentation.
- The two checks watch the same review, so the human sample sets the weekly percent and the model judge checks the judge agreement in that instrumentation.
- The two checks watch the same review, so the human sample sets the weekly population and the model judge checks the judge agreement in that instrumentation.
- The two checks watch the same review, so the human sample sets the weekly percent and the model judge checks the judge annotation in that instrumentation.