問題文
An agent's answers are accurate but the ninety-fifth percentile latency exceeds the service objective. Profiling shows most of the time is spent generating a long explanation that users rarely read. What is the most direct tuning step?
選択肢
- Add more replicas of the model server, so more requests can be served at once.
- Raise the maximum output length so the model does not have to plan its explanation around a limit, because planning around a limit is itself a source of additional generation and therefore of additional latency in the response path.
- Enable request batching, so the server processes several requests together.
- Shorten what the agent is asked to produce, so fewer tokens are generated.