問題文
A developer implements gradient accumulation over 4 micro-batches. Where should the gradient reset and the optimizer step go?
選択肢
- Reset before each micro-batch and step once after the group, so each micro-batch contributes its own gradients.
- Reset once before the group of 4, run backward for each micro-batch, then step once after the group.
- Reset after the step and never before.
- Reset and step for every micro-batch, which produces the same effective batch size because the four updates together move the parameters by the same total amount as one larger update would.