フリー問題

Databricks Certified Associate Developer for Apache Spark のフリー問題 8 / 20 問目

問題文

The same team then wants to lay out a large table by a high-cardinality user identifier so that later joins on that identifier avoid a shuffle. Which mechanism does Apache Spark provide, and what is its restriction?

選択肢

  1. Sorting the table by the identifier before writing, which lets Spark skip the exchange on later joins, and the bucket count does not matter.
  2. Caching the table in memory, which removes the need for an exchange on later joins.
  3. Bucketing with a fixed number of buckets, which is applicable only to persistent tables.
  4. Partitioning by that identifier, which works for any destination and creates one folder per distinct value so that the join can read only the matching folders on both sides without moving any data between the executors.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。