問題文
A team distributes training across a cluster because a single machine cannot hold the training data in memory. Which characterization of their situation is accurate?
選択肢
- They need model parallelism, where the layers of one single large model are split across separate devices because the parameters themselves do not fit on one machine, one shard of layers per device.
- They need data parallelism, where each worker holds a shard of the rows and the model parameters are kept in step across workers, since the limit they hit is the data size.
- They need vertical scaling only, because adding memory to one machine always resolves a data-size constraint, and the team should move the job to the largest instance available.
- They need to reduce the number of features until the data fits, because distributed training changes the resulting model, and a narrower table on one machine gives the answer they need.