問題文
A team enables automatic mixed precision on a CUDA GPU. Which pair of components does the documentation present as the typical combination for a training loop?
選択肢
- A gradient scaler alone, wrapped around the forward pass.
- An autocast context around the forward pass and the loss, together with a gradient scaler around the backward pass and the update.
- An autocast context around the whole training loop including the optimizer step, which is enough on its own because the scaler is only needed on hardware without native support for reduced precision arithmetic.
- A dtype argument passed to the optimizer.