問題文
A job needs a small lookup dictionary available inside a transformation on every executor, and also needs a running count of malformed rows visible on the driver at the end. Which pair of shared variables fits in Apache Spark?
選択肢
- A broadcast variable for both, with a second broadcast created after each partition.
- A broadcast variable for the lookup, and an accumulator for the count.
- A cached DataFrame for the lookup and a global Python variable for the count.
- An accumulator for the lookup and a broadcast variable for the count, because accumulators are the only shared variables that can be read inside a transformation and broadcast variables are the only ones the driver can inspect after the job finishes.