問題文
A deep network is built by stacking many blocks. Each block adds its own transformation to its input rather than replacing it. What does this addition primarily achieve during training?
選択肢
- It gives gradients a short path back to earlier layers, so very deep stacks remain trainable.
- It reduces the number of parameters in the network, and the added input replaces one of the two weight matrices that a block would otherwise need in order to map its input to its output dimension.
- It removes the need for normalization inside each block, since the added input already keeps the activations in range.
- It makes the network invariant to the order of the blocks, so blocks can be permuted after training without changing the output.