問題文
A team trains a model that takes an image and a sentence and learns to place matching pairs close together in one embedding space. Early in training the loss drops fast and then jumps to a very large value and stays there. The learning rate was set to the value used for a much smaller model. What is the most reasonable first action?
選択肢
- Lower the learning rate and add a warmup period so that early updates are small while the two encoders are still far from alignment.
- Increase the batch size so that each update averages over more pairs, which will smooth the gradient enough that the original learning rate becomes appropriate again without any change to the schedule itself.
- Switch the loss to a squared error between the two embeddings.
- Freeze the image encoder for the whole run so only the text side changes.