問題文
A pipeline stage recomputes the same expensive intermediate result three times because it is referenced by three downstream branches. What is the appropriate fix on a distributed cluster?
選択肢
- Write it to disk and read it back in each branch, so the graph is evaluated once and all three branches then share the same file.
- Increase the number of partitions so each branch is cheaper, since the three passes then divide into much smaller pieces of work and finish sooner.
- Materialize it once into worker memory so the three branches read the computed partitions instead of recomputing the graph.
- Gather it into the client process once and pass the resulting single object to each of the three branches, because a value that lives in the client is available to every branch without any further scheduling.