問題文
A feature is built as the mean of the target for each category. Which construction avoids leaking the target into the training features?
選択肢
- Compute the category means over the whole table once, before the split, so that the same encoding is applied consistently to every row and the training and holdout parts cannot end up with different values for the same category, which would otherwise make the holdout score hard to interpret.
- Add small random noise to the category means so that the encoding cannot be inverted to recover the target, which keeps the useful signal while masking the individual values it was built from.
- Use the category counts instead of the target means.
- Compute each row's encoding from the target values of other rows only, using out-of-fold means, and apply the encoding learned on the training part to the holdout.