フリー問題

NVIDIA-Certified Professional: Generative AI LLMs のフリー問題 15 / 20 問目

問題文

A team must distribute the matrix multiplications of multi-head attention across devices with as little communication as possible. Which partitioning follows from the structure of the computation?

選択肢

  1. Split the sequence across devices for the attention product; each position's output depends only on the positions that come before it in the sequence.
  2. Assign whole heads to devices, since each head's projections and its attention product are independent until the outputs are combined.
  3. Split the batch across devices and keep all heads together, the batch dimension being the one that carries no dependency between its entries.
  4. Split each head's key dimension across devices; the key dimension is the axis the products are summed over, so dividing it gives every device an equal share of the arithmetic with a single reduction at the end.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。