問題文
A model must fit into a smaller memory budget and also run faster on hardware that accelerates dense matrix multiplication. Which pruning approach matches both goals?
選択肢
- Structured pruning that removes whole units such as attention heads or feed-forward channels.
- Reducing the batch size at inference time.
- Unstructured magnitude pruning applied uniformly across all layers, which removes the least important individual weights and gives the best accuracy for a given reduction in the total number of nonzero parameters.
- Quantization of the weights to a lower-precision integer format, which shrinks the matrices the hardware multiplies.