問題文
A team wants to know whether a new prompt template raises the resolution rate of a support assistant. They plan to serve the new template to users who contact support in the morning and the old one to users who contact in the afternoon. Why is that design weak?
選択肢
- The sample size will be too small in the morning group, and once both groups reach several thousand conversations the time-of-day split becomes valid on its own.
- Prompt templates cannot be compared with an experiment at all, and the same template produces different answers on repeated requests to the same model.
- The two groups differ in ways other than the template, so any difference in outcome cannot be attributed to the change.
- The old template will be at a disadvantage: it has already been optimized, and its remaining headroom is smaller than the new one's.