フリー問題

Databricks Certified Machine Learning Professional のフリー問題 5 / 20 問目

問題文

An insurance company trains a claim severity model on a Delta table with 620 million rows. The table has three numeric columns and two string columns holding the policy region and the claim category. The team wants the whole preparation to run on the cluster and to keep working as the table grows each quarter. Which preparation should they build?

選択肢

  1. Index each string column, expand the indexes into indicator columns, and gather every numeric and indicator column into one vector column with VectorAssembler, since each of those steps runs as a Spark job.
  2. Collect the table to the driver as a pandas DataFrame, build indicator columns there, and hand the result to a single-node estimator, which caps the work at what one machine holds.
  3. Write a Python function that encodes one row at a time and register it as a scalar user-defined function applied over the two string columns; the encoding happens one row at a time inside the generated query plan.
  4. Cast the two string columns to double so that the estimator can consume them directly, and skip the assembling step, which keeps the stage count low.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。