問題文
A speech dataset is assembled from three sources: call-center recordings at 8 kHz, podcast audio at 44.1 kHz, and field recordings at 16 kHz. The pretrained speech model the team plans to use expects 16 kHz input. What should the preprocessing stage do?
選択肢
- Downsample every clip to 8 kHz so that all sources share the lowest common rate available in the collection, and treat one shared rate as worth more than the rate the model was trained on.
- Resample every clip to 16 kHz so that the model receives the sampling rate it was trained on, and record which clips were upsampled.
- Keep each clip at its original rate and let the model infer the rate from the waveform, because the length of the array and the amplitude envelope together carry enough information for the feature extractor to normalize the input on its own.
- Convert all clips to a compressed format so the file sizes match.