問題文
An analyst has a 40 GB directory of Parquet files but needs only four of its sixty columns for a report. Which reading strategy keeps the transfer to GPU memory smallest?
選択肢
- Convert the directory to a single comma-separated file first and then read the four columns from it, so all sixty fields are read first.
- Read the whole directory and rely on the memory manager to spill, so the shortfall turns into extra time while the decode stays as large.
- Pass the list of four column names to the Parquet reader so that only those columns are decoded and materialized.
- Read every column into a dataframe first and then drop the unwanted ones, because the reader has to parse the whole file layout anyway and dropping afterwards releases the same amount of memory without complicating the call.