問題文
In a training batch of image-and-caption pairs, a contrastive objective computes similarity between every image and every caption in the batch. What does this objective ask the model to do?
選択肢
- Classify each image into one of a fixed set of categories, and treat each caption in the batch as the name of one of those categories.
- Minimize the absolute distance between each image embedding and its own caption embedding, without any reference to the other items in the batch, so that matching pairs eventually occupy exactly the same point in the shared embedding space.
- Raise the similarity of each image with its own caption relative to its similarity with the other captions in the batch.
- Predict the next word of each caption given the image.