問題文
A developer needs a custom scoring function applied to a numeric column of a large DataFrame, and wants the function to receive many values at a time rather than one value per call. Which mechanism fits in PySpark?
選択肢
- A plain Python user-defined function, because the engine automatically groups the rows of a partition into a single call when the function is registered with a return type that Spark can vectorize on its behalf.
- A grouped map operation, applied per group of rows.
- A pandas UDF that takes a series and returns a series of the same length.
- A broadcast variable holding the function, which each executor then applies row by row.