フリー問題

NVIDIA-Certified Professional: Generative AI LLMs のフリー問題 8 / 20 問目

問題文

A service must raise its supported context from 4,000 to 32,000 tokens on the same hardware, keep the first-token latency acceptable, and avoid retraining. Which combination of measures fits those constraints?

選択肢

  1. Raise the batch size so the longer inputs are processed together.
  2. Disable the cache of keys and values to save memory.
  3. Adopt a serving stack with paged storage for the keys and values and a tiled attention implementation, and lower the precision of the stored keys and values, validating quality on long inputs.
  4. Retrain the model with a windowed attention pattern so that the memory for the stored keys and values stops growing with the input length; a fixed window bounds the memory regardless of how long the context becomes, and this is the only measure that removes the growth entirely.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。