問題文
An analyst runs the same aggregation on 3 million rows with a processor-only library and with an accelerated library, and finds the accelerated run is slightly slower. What is the most likely explanation, and what should they do?
選択肢
- The measurement is invalid because the two libraries produce different results for the same aggregation, so no timing comparison is meaningful.
- At this size the fixed costs of moving the data to the device and launching the work dominate the saving, so they should measure again at the size the production job actually has before deciding.
- The device needs a larger memory pool.
- The accelerated library is slower for aggregations in general, because grouping requires the rows to be brought together and that step cannot be parallelized at all, so aggregation-heavy jobs should always be left on the processor no matter how large the input becomes.