問題文
A team trains a model that takes audio and video together. The audio branch converges quickly while the video branch barely improves, and the combined loss plateaus. What is a reasonable interpretation?
選択肢
- The video branch has too many parameters relative to the audio branch, and reducing its capacity will let it fit the remaining signal because a smaller network converges faster on the same amount of data than a larger one does.
- The two modalities are not synchronized in time, so the video branch cannot find a signal that matches the audio it is paired with.
- The easier modality is satisfying the objective on its own, so the gradient reaching the harder branch is small.
- The learning rate is too low for both branches, so neither branch is moving and the plateau is a step-size problem rather than an imbalance.