問題文
A large fact DataFrame must be enriched with attributes from a small reference DataFrame, and every fact row must survive even when no reference row matches. Which join type meets that requirement in PySpark?
選択肢
- A left semi join, which keeps the fact rows and adds the reference attributes.
- An inner join, because the reference DataFrame is a complete dimension by construction and therefore every fact row is guaranteed to find a match, which makes the outer variant an unnecessary cost.
- A full outer join, which is the only way to keep unmatched rows on either side, so unmatched reference rows appear as well and have to be filtered out afterwards.
- A left outer join with the fact DataFrame on the left, which keeps unmatched fact rows with null attributes.