問題文
Before launching the comparison, a reviewer insists that the team write down which single measure will decide the outcome and how long the run will last. What problem does that practice prevent?
選択肢
- Introducing a second change halfway through the run, which a fixed duration and a single measure both rule out.
- Choosing whichever measure happens to look favorable after the fact, which makes a chance difference look like a real effect.
- Running out of traffic before the experiment concludes, which is the usual reason that a comparison of two prompt templates ends without a usable answer.
- Assigning users to groups unevenly, and a fixed duration guarantees that both groups receive the same number of conversations over the run.