フリー問題

Databricks Certified Associate Developer for Apache Spark のフリー問題 20 / 20 問目

問題文

A developer needs a custom scoring function applied to a numeric column of a large DataFrame, and wants the function to receive many values at a time rather than one value per call. Which mechanism fits in PySpark?

選択肢

  1. A plain Python user-defined function, because the engine automatically groups the rows of a partition into a single call when the function is registered with a return type that Spark can vectorize on its behalf.
  2. A grouped map operation, applied per group of rows.
  3. A pandas UDF that takes a series and returns a series of the same length.
  4. A broadcast variable holding the function, which each executor then applies row by row.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。