問題文
A team reviews a collection of 500,000 image-and-caption pairs before training. They have time to inspect only a small portion by hand. Which sampling approach is most likely to surface data problems?
選択肢
- A mix of a uniform random sample and deliberately chosen extremes, such as the shortest captions and the images with unusual dimensions.
- A uniform random sample only, because it is the unbiased estimator of the collection and any deliberate selection would give a distorted picture of how frequent each kind of problem is across the whole dataset.
- The first 1,000 items in the file: ingestion order is arbitrary.
- The items with the highest similarity between image and caption.