フリー問題

NVIDIA-Certified Associate: Generative AI Multimodal のフリー問題 8 / 20 問目

問題文

A team trains a model that takes audio and video together. The audio branch converges quickly while the video branch barely improves, and the combined loss plateaus. What is a reasonable interpretation?

選択肢

  1. The video branch has too many parameters relative to the audio branch, and reducing its capacity will let it fit the remaining signal because a smaller network converges faster on the same amount of data than a larger one does.
  2. The two modalities are not synchronized in time, so the video branch cannot find a signal that matches the audio it is paired with.
  3. The easier modality is satisfying the objective on its own, so the gradient reaching the harder branch is small.
  4. The learning rate is too low for both branches, so neither branch is moving and the plateau is a step-size problem rather than an imbalance.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。