フリー問題

Databricks Certified Associate Developer for Apache Spark のフリー問題 14 / 20 問目

問題文

A production job reads JSON files whose fields are known and fixed. The team wants to avoid the extra pass over the data that inference costs and to fail fast if a field type changes. What should they do in PySpark?

選択肢

  1. Sample a single file for inference and assume the rest matches, and the cost of the extra pass is halved if the files are alike.
  2. Read the files as text and parse them with a user-defined function, and no extra pass over the data is needed.
  3. Enable inference and cache the inferred schema in the session so that later runs reuse it, because the inference pass then happens only once for the whole application and the types are still checked against the cached definition on every run.
  4. Supply an explicit schema to the reader instead of relying on inference.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。