問題文
A nightly extract from a busy operational database must finish in an hour without hurting the source system. Which pair of decisions fits in Apache Spark?
選択肢
- Read with a very large fetch size so fewer round trips are needed and no partitioning is necessary.
- Read the whole table in one partition to minimize the load, and accept a longer run time.
- Use as many read partitions as the cluster has slots so the extract finishes as fast as possible, and rely on the source database to queue the connections it cannot serve immediately so that the extract never has to wait for capacity to become available on the remote side.
- Use a moderate number of read partitions with bounds matched to the real key range, and schedule the extract in the source system's quiet window.