問題文
A pipeline imputes missing numeric entries with the column mean. Where must that mean be computed to avoid leakage?
選択肢
- From the training subset only, with the same stored value then applied to the evaluation subset and to production records.
- From each subset separately, using its own mean, so every subset is filled with a value that matches the distribution of that same subset.
- From the production records at inference time, so the filled value tracks the traffic arriving at that moment and follows it as it drifts.
- From the combined training and evaluation subsets, because using more rows gives a more accurate estimate of the column mean and a more accurate imputation can only help the model generalize better.