フリー問題

NVIDIA-Certified Associate: Generative AI Multimodal のフリー問題 6 / 20 問目

問題文

In a training batch of image-and-caption pairs, a contrastive objective computes similarity between every image and every caption in the batch. What does this objective ask the model to do?

選択肢

  1. Classify each image into one of a fixed set of categories, and treat each caption in the batch as the name of one of those categories.
  2. Minimize the absolute distance between each image embedding and its own caption embedding, without any reference to the other items in the batch, so that matching pairs eventually occupy exactly the same point in the shared embedding space.
  3. Raise the similarity of each image with its own caption relative to its similarity with the other captions in the batch.
  4. Predict the next word of each caption given the image.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。