問題文
A pipeline is being written for a shared landing zone where several teams may write the same dataset. The requirement is that a second writer must leave the existing data completely alone and must not fail the job. Which save mode fits in Apache Spark?
選択肢
- Append mode, because adding nothing is equivalent to leaving the data alone, so a second run repeats the rows and the table grows.
- Ignore mode, which skips the write when data already exists and leaves it unchanged.
- Overwrite mode with a filter that excludes rows already present at the destination, and the earlier rows are left in place.
- The default mode, because raising an exception is the safest way to protect data that another team has already written, and the surrounding scheduler can be configured to treat that particular exception as a successful completion.