問題文
A dashboard shows the number of distinct visitors per hour over 50 million rows and refreshes every hour. A 2 percent error in the distinct count is acceptable and the refresh must be fast. Which aggregate fits, and why?
選択肢
- A collect of the visitor column on the driver followed by a Python set.
- A plain row count with a filter, because distinct visitors can be derived from the row count.
- The exact distinct count, because the engine keeps a sorted structure of the values that makes the exact answer as cheap as the approximation once the data is large enough for the sorting cost to be amortized across the partitions.
- The approximate distinct count, because a probabilistic sketch avoids the cost of collecting every distinct value.