問題文
A retailer must train a separate forecasting model for each of 4,000 stores. Each store has at most 90,000 rows and the same model type. The team wants all 4,000 fits to run in parallel on one cluster. Which mechanism fits?
選択肢
- Group the DataFrame by store and apply a pandas function that receives each group as a pandas DataFrame, fits the model, and returns the results with a declared output schema, since the schema lets Spark plan the 4,000 tasks.
- Register the fitting code as a scalar user-defined function and call it once per row of the original table, once for each event row in the table.
- Train one distributed SparkML model on all stores at once and add the store identifier as an indicator feature, giving one shared set of coefficients for all stores.
- Loop over the 4,000 store identifiers in the driver, filtering the DataFrame and fitting one model per iteration; the driver issues 4,000 sequential jobs and keeps each result in memory.