問題文
An architect notices that a model which occupies a certain amount of accelerator memory when answering requests does not fit on the same accelerator during training, even with a batch of one. A junior engineer is surprised. What explains the difference?
選択肢
- Training stores the whole dataset in accelerator memory at the start of the run so that every step reads its samples fast.
- Training keeps a copy of the weights for every accelerator in the job, even on a single-accelerator run.
- Serving compresses the weights to a smaller numeric format, which is why it needs less memory than training.
- Training also holds gradients, optimizer state, and the activations needed for the backward pass, none of which serving keeps.