問題文
A training data set consists of tens of millions of small files. Why can the storage system become the bottleneck even when its published throughput figure looks sufficient?
選択肢
- Because a file smaller than the network's maximum transmission unit cannot be transferred at all and has to be padded by the client before it is sent, which is what consumes the extra time.
- Because small files are always stored on a slower tier than large files, regardless of how the system is configured, so the fix is to place the data set on the faster tier by hand at build time.
- Because the published throughput figure is measured with small files, so it understates what the system can do with a data set of this shape, and the measured rate comes out above the published one.
- Because each small file costs a lookup as well as a transfer, so the rate of metadata operations can run out long before the data transfer rate does.