問題文
A team must distribute the matrix multiplications of multi-head attention across devices with as little communication as possible. Which partitioning follows from the structure of the computation?
選択肢
- Split the sequence across devices for the attention product; each position's output depends only on the positions that come before it in the sequence.
- Assign whole heads to devices, since each head's projections and its attention product are independent until the outputs are combined.
- Split the batch across devices and keep all heads together, the batch dimension being the one that carries no dependency between its entries.
- Split each head's key dimension across devices; the key dimension is the axis the products are summed over, so dividing it gives every device an equal share of the arithmetic with a single reduction at the end.