問題文
A nightly job reads 2 TB of raw event files, filters to one region, joins a small lookup table, aggregates by day, and writes the result. Which ordering of those steps moves the least data through the accelerated stack?
選択肢
- Join the lookup table first so that the region name is available as a readable label, then filter on that label, then aggregate, because filtering on a human-readable name is easier to review and the join itself is cheap when one side of it is small.
- Aggregate by day first to shrink the row count, then filter to the region, then join the lookup table, and the daily totals are far smaller than the raw events and every later step then works on that reduced table.
- Read everything, then let the query optimizer reorder the steps, and the optimizer sees the whole plan and will push the region filter down into the read for the team automatically.
- Filter to the region and select only the needed columns while reading, then join the lookup table, then aggregate.