フリー問題

NVIDIA-Certified Associate: Generative AI LLMs のフリー問題 7 / 20 問目

問題文

A generation endpoint serves many concurrent users. The team wants to raise the number of requests completed per minute on the same hardware. Which change targets that measure most directly?

選択肢

  1. Return the tokens to the client as they are produced, which lets the server begin the next request sooner and therefore raises the number completed per minute.
  2. Increase the maximum output length so that fewer requests are truncated and have to be retried by the client, which is the usual cause of low completion counts.
  3. Group concurrent requests so that the accelerator processes several of them together in each step rather than one at a time.
  4. Add more parallel attention computations per layer, so each step of the accelerator covers more requests at once.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。