問題文
A multi-process data loader inside a training container fails with an error about insufficient shared memory as soon as the number of worker processes is raised. The same code runs on the host. What is the usual cause and remedy?
選択肢
- The container needs more devices, since each worker process requires one and a worker that finds none falls back to a buffer far smaller than it asked for.
- The container's shared memory allocation defaults to a small size, and the container has to be started with a larger one.
- The image lacks the accelerator libraries, and the loader reports that as a shared memory error once the first worker tries to map its buffers.
- The host's kernel parameters for shared memory are too low and have to be raised on the node, since a container inherits those limits directly from the host and cannot be given a different value of its own.