問題文
A grouped pandas function returns a pandas DataFrame with columns named store, forecast date, and predicted units. The job fails before any group is processed. What is the most likely cause?
選択肢
- The output schema was not declared, so Spark cannot know the types of the returned columns before running the function, since the plan is built before any group runs.
- The number of returned rows must equal the number of input rows in the group, so a 14-row forecast from a 90,000-row group is rejected, and the function must pad its output to match.
- The grouping column must be excluded from the returned columns, so including the store column raises an error, and dropping store from the returned frame lets groups run.
- The function returned a pandas DataFrame instead of a Spark DataFrame, which is never allowed, and the group API hands back a Spark frame instead.