問題文
On a freshly configured queue-based cluster, users report that their jobs land on accelerated nodes and run, but every job sees no accelerators at all. The node definitions and the per-node resource file both list the devices correctly. What explains the observation?
選択肢
- The accelerators are visible only to the first job step of each allocation, so subsequent steps see none of the devices the job was granted at submit time.
- The jobs do not request the generic resource at submit time, and a job is never allocated one unless it asks for it.
- The devices must be listed in the partition definition as well, and a partition that omits them hands out plain cores to every job it accepts.
- The devices are declared but the node has not been returned to service after the declaration, so the scheduler is still working from the resource counts it cached when the node first registered with the controller.