問題文
Two prompt templates are being compared on the same 200 questions. The team runs each template once with sampling enabled and different random seeds, then declares a winner from the single pair of scores. What is the flaw?
選択肢
- With sampling enabled, one run per template mixes the effect of the template with the randomness of decoding, so the difference may not be reproducible.
- The scores cannot be compared, and the two templates produce answers of different lengths, and length alone decides the score.
- Two hundred questions is too few for any comparison, and no design can produce a usable answer until the evaluation set holds several thousand items.
- Comparing templates requires that both be evaluated on different question sets so that neither template can benefit from questions it has already been tuned against.