フリー問題

NVIDIA-Certified Associate: Generative AI LLMs のフリー問題 10 / 20 問目

問題文

Users of a chat feature complain that nothing appears on screen for several seconds when the answer is long, even though the total time is acceptable. The team wants to improve the perceived responsiveness without changing the model. What should they do?

選択肢

  1. Cache the whole response for identical requests, which helps only when the same question is asked twice in a row.
  2. Increase the maximum output length so that the response is generated in one larger step, which lets the server return everything at once rather than in pieces.
  3. Send the generated tokens to the client as they are produced instead of waiting for the full response.
  4. Lower the decoding randomness, and a more deterministic decode finishes each token faster and therefore shortens the delay before the first character appears.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。