問題文
A production job reads JSON files whose fields are known and fixed. The team wants to avoid the extra pass over the data that inference costs and to fail fast if a field type changes. What should they do in PySpark?
選択肢
- Sample a single file for inference and assume the rest matches, and the cost of the extra pass is halved if the files are alike.
- Read the files as text and parse them with a user-defined function, and no extra pass over the data is needed.
- Enable inference and cache the inferred schema in the session so that later runs reuse it, because the inference pass then happens only once for the whole application and the types are still checked against the cached definition on every run.
- Supply an explicit schema to the reader instead of relying on inference.