問題文
A team is building similarity search over 50,000 knowledge articles. Some articles were embedded last year with one model and the rest this year with a different one. Users report that results are erratic. What explains this?
選択肢
- Similarity scores are only valid for text under 512 characters, and the longer product descriptions fall outside the supported range and score at random.
- The two models produce vectors of different lengths, so the comparison silently pads the shorter vectors with zeros and the padding dominates the resulting score.
- Vectors produced by different embedding models are not comparable, so similarity computed across the two sets is meaningless.
- The cosine similarity function requires normalized vectors, and one of the two model outputs is left unnormalized before the comparison.