問題文
For multi-GPU training on a single machine, which approach does the documentation recommend, and why?
選択肢
- The single-process wrapper that replicates the model across devices, because keeping everything in one process removes the need for inter-process communication and therefore scales better on a single machine.
- Manual splitting of the batch with explicit copies to each device.
- Distributed data parallel training with one process per device, because it avoids the bottleneck of a single process driving all devices.
- Running independent processes without any gradient synchronization.