問題文
A DataFrame is reused by three separate actions, so a developer calls persist on it without passing any argument. Which storage behavior does Apache Spark apply for a DataFrame in that case?
選択肢
- Partitions are kept in memory only, and any partition that does not fit is dropped and recomputed from the source the next time it is needed by a later action in the application.
- Nothing is stored until an action runs, and then the data is copied to the driver, so the three actions each read the source files again and no blocks are retained.
- Partitions are written to disk only, to keep the memory free for shuffles.
- Partitions are kept in memory, and partitions that do not fit are written to disk rather than being discarded.