問題文
Checkpoint writes have become much slower while the training steps themselves are unchanged. How can the storage system be distinguished from the path to it?
選択肢
- Compare the accelerator utilization during the checkpoint with the utilization during the training steps, since a difference identifies the storage system, and the size of that difference gives the scale of the delay.
- Compare a write issued from a node on a different part of the fabric with one issued from the affected nodes, and look at the counters on the links each of them uses.
- Increase the checkpoint interval so that fewer writes are issued, which removes the symptom and confirms that the storage system rather than the fabric was the limiting factor in the original configuration.
- Restart the storage system, since a slow write path always recovers if the storage system is the cause, and the writes return to their previous rate as soon as the service comes back.