問題文
An events table receives the same event more than once when a producer retries. Each copy carries the same event identifier but a different ingestion timestamp. A data scientist needs one row per event, keeping the earliest ingestion. Which approach produces that result?
選択肢
- Select the distinct rows of the table, which the producer writes with an ingestion timestamp, because distinct removes repeated rows and the retries are repeats.
- Group by the event identifier and take the count of rows, because the count tells which events were retried by the producer and those can then be excluded.
- Take the minimum of the ingestion timestamp for each event identifier, because the minimum identifies the earliest copy and every other column of that same row follows from it.
- Number the rows within each event identifier by ascending ingestion timestamp, then keep only the rows numbered one.