問題文
A summarization feature is scored by comparing each generated summary against one human-written reference and counting shared word sequences. Editors report that the newer summaries read better, yet the score is unchanged. What does that combination indicate?
選択肢
- The metric needs a larger evaluation set to detect the change.
- The editors must be mistaken; a metric computed against a human-written reference is by construction a measure of how close the output is to what a human would have written and cannot miss a genuine improvement in readability.
- The generated summaries are too short for the metric to apply, so the count of shared word sequences stays flat until the summaries grow longer.
- The metric rewards surface overlap with one reference, so an improvement expressed in different wording is invisible to it.