問題文
An analyst is asked to produce the number of distinct visitor identifiers in a table of 4 billion page views. The figure appears on an exploratory dashboard and a small error is acceptable if the query returns quickly. Which approach matches the requirement?
選択肢
- Count all rows and divide by the average number of page views per visitor taken from last month's report.
- Sample one percent of the rows at random, count the distinct identifiers in the sample, and multiply the result by one hundred.
- Use approx_count_distinct on the identifier column for that figure, which trades a small bounded error for far less work than an exact distinct count.
- Use COUNT with DISTINCT on the identifier column, because exact answers are always preferable and the optimizer will approximate it automatically when the table is large.