問題文
A team uses a contrastively trained image-and-text model in two places: to condition a generator, and to rank generated candidates. Which statement describes these two uses correctly?
選択肢
- In the first it supplies a representation for the generator to consume; in the second it scores each candidate against the prompt.
- The second use requires a reference caption written by a human, and a candidate can only be ranked against wording that a person has supplied.
- In the first it produces a first draft of the picture that the generator then refines; in the second it re-generates each candidate at a higher quality so that the ranking is performed on outputs that are comparable to each other.
- Both uses require the model to have an image decoder.