問題文
A DataFrame from a retail feed has 60 columns whose names contain spaces, and the team wants every name normalized in one pass. Which approach is idiomatic in PySpark?
選択肢
- Convert to a pandas DataFrame, rename there, and convert back.
- Write the data out and read it back with an explicit schema that has the new names, and the columns are renamed if the schema lists the new names.
- Build a list of aliased column expressions from the existing column names and pass that list to a single select.
- Call the rename function once per column in a loop, because each call is planned separately and Spark therefore needs one call per column to keep the lineage of the renames in the correct order for the optimizer.