問題文
A team wants to find, for a given question, the 20 internal documents most likely to contain the answer, out of 900,000 documents. Which use of vector embeddings is correct?
選択肢
- Convert only the question into a vector and compare it with the raw text of the documents, because the documents can be matched against the vector as ordinary strings.
- Convert the documents and the question into vectors, then rank the documents by the similarity between their vectors and the question vector.
- Fine-tune a model on the 900,000 documents first, because similarity search requires a tuned model.
- Ask the embedding function to generate the answer directly, because embeddings return the text in each internal document that best matches the question.