問題文
A cluster has four nodes of eight devices each. Devices inside a node are joined by a fast interconnect; the nodes are joined by a much slower network. A model needs four devices' worth of memory for one layer's weights. Which assignment of the three parallelism axes fits the topology?
選択肢
- Split each layer's tensors across nodes so that the eight devices inside a node can each hold a different layer group, because putting the layer groups inside a node keeps the activations that pass between consecutive groups on the fast interconnect where the largest volume of traffic occurs.
- Replicate the model on every device and average gradients; replication needs no coordination inside a layer and the slower network only carries one collective per step.
- Split each layer's tensors within a node, place consecutive layer groups across nodes, and replicate the whole arrangement for the remaining capacity.
- Place consecutive layer groups within a node and split tensors across nodes, so the fast interconnect carries the activations between the consecutive groups.