問題文
An intermittent problem affects a 64 node training job but cannot be reproduced on any single node. Which approach isolates the cause most efficiently?
選択肢
- Reduce the size of the job to a single node, confirm that it works, and then declare the infrastructure healthy on the grounds that the smallest reproducible unit shows no fault and the problem must therefore lie in the application.
- Run the same job on halves of the cluster in turn, then keep halving whichever part still shows the problem.
- Increase the logging level on every device at once and wait for the problem to appear again, so the full record of the event is available from every device the moment it happens.
- Replace the switches one at a time until the problem disappears, starting from the spine layer, and the last replacement made before the problem clears identifies the faulty device.