問題文
A distributed training job is spread across 16 nodes. What does that structure imply for how node level failures must be handled?
選択肢
- Because the remaining 15 nodes automatically absorb the failed node's share, the job continues at 15 out of 16 of its original speed with no further action.
- Because each node works on an independent copy of the problem, a failure affects only the results produced by that node and the rest of the run is unaffected.
- Because every node participates in each step, losing one node stops the job, so the design must rely on periodic checkpoints to bound how much work is lost.
- Because the job is stateless, a failed node can be replaced at any time by an idle node from the shared pool, and the run resumes from exactly the point of failure without any prior arrangement being necessary.