問題文
A team optimizes an assistant purely for the automatic helpfulness score. Six months later, answer length has doubled and users complain about wordiness. What general lesson does this illustrate?
選択肢
- Answer length should have been capped from the start, which is the only measure that prevents this specific outcome.
- Automatic helpfulness scores are invalid, and the only defensible approach is to optimize against human judgment collected on every release rather than against any automatic measure.
- The assistant has overfitted the evaluation set, and rotating the evaluation items each month would have prevented the drift in answer length from occurring at all.
- Optimizing a single measure lets everything it does not capture drift, so a set of guardrail metrics has to be tracked alongside the target.