問題文
A telecom company trains a churn model on last quarter's data. The plan code column had 14 values then, but marketing launches two new plans next month. The scoring job must not fail when a new plan code appears. What should the team configure on the indexing stage?
選択肢
- Replace the indexing stage with a hash of the plan code string, because hashing has no notion of unseen values, so every string maps.
- Set the handling of unseen values so that rows with an unknown plan code are assigned to a dedicated extra index rather than raising an error.
- Retrain the indexing stage on the scoring data before each run, so that the mapping always covers the values present at that time, including the new plan codes.
- Filter out the rows that carry unknown plan codes before scoring them at all, because a model cannot produce a valid prediction for an unseen category in any case.