問題文
A service must raise its supported context from 4,000 to 32,000 tokens on the same hardware, keep the first-token latency acceptable, and avoid retraining. Which combination of measures fits those constraints?
選択肢
- Raise the batch size so the longer inputs are processed together.
- Disable the cache of keys and values to save memory.
- Adopt a serving stack with paged storage for the keys and values and a tiled attention implementation, and lower the precision of the stored keys and values, validating quality on long inputs.
- Retrain the model with a windowed attention pattern so that the memory for the stored keys and values stops growing with the input length; a fixed window bounds the memory regardless of how long the context becomes, and this is the only measure that removes the growth entirely.