問題文
A generation endpoint serves many concurrent users. The team wants to raise the number of requests completed per minute on the same hardware. Which change targets that measure most directly?
選択肢
- Return the tokens to the client as they are produced, which lets the server begin the next request sooner and therefore raises the number completed per minute.
- Increase the maximum output length so that fewer requests are truncated and have to be retried by the client, which is the usual cause of low completion counts.
- Group concurrent requests so that the accelerator processes several of them together in each step rather than one at a time.
- Add more parallel attention computations per layer, so each step of the accelerator covers more requests at once.