問題文
While reviewing a transformer implementation, an engineer finds a step that adds a position-dependent signal to the token representations before the first attention layer, and wonders whether deleting it would simply simplify the code. What would break?
選択肢
- Sequences longer than the training length would fail with an error, while shorter ones would keep working exactly as before because they fit inside the allocated position table.
- The model would still know the order, and the tokenizer emits tokens in order, but the attention weights would become harder to interpret during debugging.
- Attention treats the input as an unordered collection, so without a position signal the model could not tell two sentences with the same words in a different order apart.
- The loss would stop decreasing entirely because gradients cannot flow through an attention layer that has no position term, so the run stalls on the first pass.