問題文
Before launch, a team wants to know how many concurrent users their agent can serve while keeping the ninety-fifth percentile latency under the objective. What kind of measurement answers this?
選択肢
- The token count of a typical request multiplied by the expected number of users.
- A single-request timing measurement, repeated one hundred times.
- A load test that drives increasing concurrency against the workflow and records latency and throughput at each level.
- A measurement of the model server's peak throughput in isolation, because the model call is the slowest part of the workflow and the workflow's own capacity is therefore determined entirely by the throughput that the server can sustain.