問題文
A team wants to let users supply a reference photograph so generated images match its subject. The pipeline currently accepts only text. What is the conceptually straightforward extension?
選択肢
- Generate a caption for the reference photograph and use it as the prompt, so the subject reaches the pipeline through text and nothing else has to change.
- Concatenate the reference photograph to the initial noise tensor along the channel dimension, because the network then sees the reference at every step and will reproduce its subject in the output without any change to how the condition is supplied.
- Train a new model from scratch on the reference images.
- Encode the reference image into the same aligned space the text uses, and pass that representation as the condition.