フリー問題

Databricks Certified Machine Learning Professional のフリー問題 15 / 20 問目

問題文

A team writes an integration test for their machine learning pipelines. They want the test to finish in a few minutes on every pull request. What is the appropriate handling of the data the test uses?

選択肢

  1. Use the full production tables but limit the cluster size so that the run stays within the time budget, so the same tables are read at a smaller degree of parallelism.
  2. Generate synthetic data with random values on each run, because the test only needs to verify that the code executes without raising, and a run that finishes is the signal the team wants.
  3. Use a small fixed dataset that keeps the same schema and the same edge cases as production, stored where the test can read it reproducibly, since a fixed input makes a failure readable.
  4. Sample a fresh random slice of the production data on every single run so that the test always reflects the conditions the pipelines are currently seeing in production, one slice per pull request.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。