問題文
A support assistant must answer in under two seconds, but the current pipeline retrieves ten passages, reranks them, and uses a large model. Which change gives the most latency improvement for the least quality loss?
選択肢
- Remove the retrieval step entirely and rely on the model's own knowledge, since the two-second target is a hard requirement and retrieval is the single largest contributor to the delay
- Raise the temperature so the model commits to an answer faster instead of weighing several equally likely continuations at each step
- Reduce the number of retrieved passages, after measuring how many of them are actually needed for the answers to stay correct, so the input shrinks without the quality falling
- Increase the max output tokens so the model finishes sooner, since a generation that is not constrained by the limit does not need to compress its wording as it goes