問題文
An engineer migrating a summarization service from an older recurrent design to a transformer design reports that training throughput improved far more than the parameter count alone would suggest. Which property of self-attention best explains that?
選択肢
- It computes the relationships among all positions of a sequence at once, so the work can be spread across the accelerator instead of waiting for the previous position.
- It compresses the input down to a fixed number of positions before the first layer of the network runs at all, so a long input costs the same as a short one.
- It replaces floating-point arithmetic with integer arithmetic throughout the network, which is what makes the accelerator faster on the same hardware.
- It removes the need to store any intermediate values during the forward pass, so the same accelerator can hold a larger batch and the time per pass falls with it.