フリー問題

Implementing Cisco Data Center AI Infrastructure のフリー問題 8 / 20 問目

問題文

A distributed training job is spread across 16 nodes. What does that structure imply for how node level failures must be handled?

選択肢

  1. Because the remaining 15 nodes automatically absorb the failed node's share, the job continues at 15 out of 16 of its original speed with no further action.
  2. Because each node works on an independent copy of the problem, a failure affects only the results produced by that node and the rest of the run is unaffected.
  3. Because every node participates in each step, losing one node stops the job, so the design must rely on periodic checkpoints to bound how much work is lost.
  4. Because the job is stateless, a failed node can be replaced at any time by an idle node from the shared pool, and the run resumes from exactly the point of failure without any prior arrangement being necessary.

解答・解説を確認するには

正解と解説の確認、回答の記録には無料登録が必要です。登録すると演習モードでフリー問題に回答し、正誤と解説をその場で確認できます。