問題文
For a 400 MB table on a laptop with one device, which library choice is defensible, and on what grounds?
選択肢
- A distributed cluster with one worker, because that combination has no overhead at all compared with a plain single-process run, and the same script then scales to a larger cluster whenever the input grows.
- A distributed cluster, because starting with the distributed form from the beginning means the code will not need to be rewritten when the data grows, and the overhead of the scheduler is negligible compared with the cost of a migration later on.
- Only the processor library, because accelerated libraries require at least a few gigabytes of input, and a table of this size fits comfortably in host memory where the processor library works through it in one pass.
- Either a single-process library or the accelerated single-device library; a distributed cluster adds scheduling and communication cost that this size does not repay.