フリー問題

Databricks Certified Associate Developer for Apache Spark のフリー問題 12 / 20 問目

問題文

A DataFrame from a retail feed has 60 columns whose names contain spaces, and the team wants every name normalized in one pass. Which approach is idiomatic in PySpark?

選択肢

  1. Convert to a pandas DataFrame, rename there, and convert back.
  2. Write the data out and read it back with an explicit schema that has the new names, and the columns are renamed if the schema lists the new names.
  3. Build a list of aliased column expressions from the existing column names and pass that list to a single select.
  4. Call the rename function once per column in a loop, because each call is planned separately and Spark therefore needs one call per column to keep the lineage of the renames in the correct order for the optimizer.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。