問題文
A team needs the latest row per key from a stream of updates, keeping all the columns, before merging into a target. Which approach is both correct and expressible in the DataFrame API?
選択肢
- Sort the whole DataFrame by the timestamp and then take the first row, because a sort makes the first row the most recent one for every key at the same time.
- Remove duplicates continuously on the key column and on its plain column cast, which keeps the most recent row.
- Add a rank over a window partitioned by the key and ordered by the timestamp, then keep the rows whose rank is one.
- Group by the key and by its type cast of a query in it and take the maximum timestamp, which yields the latest row with all its columns.