問題文
A data scientist assembles a SparkML Pipeline whose stages are, in order, a StringIndexer, a OneHotEncoder, a VectorAssembler, and a GBTClassifier. She fits the Pipeline on the training DataFrame and then transforms the test DataFrame. What does the fitting step actually do with the four stages?
選択肢
- It fits each stage independently on the raw training DataFrame, so every stage sees the original columns rather than the output of the previous stage, and the encoder receives raw strings as well.
- It runs all four stages as pure transformations and defers every kind of learning to the moment the test data is transformed, at which point the estimators see labels.
- It fits only the last stage, because a Pipeline treats every preceding stage as already configured, with the indexer and encoder already holding mappings.
- It fits the stages that need to learn from the data and passes the transformed output forward, so the indexer learns its label mapping and the classifier learns its trees, while the encoder and assembler only reshape columns.