問題文
A developer wants to apply a Python side effect, such as sending a notification, once per row of a large DataFrame. Which approach is designed for that in Apache Spark?
選択肢
- A broadcast variable holding the notification client.
- The per-partition iteration action, which runs the function on the executors and lets a connection be opened once per partition.
- A collect to the driver followed by a Python loop, which is the recommended pattern because the side effect then happens exactly once per row even when a task is retried on another executor after a failure.
- A user-defined function, because side effects belong in an expression.