A task orchestrator has been configured to run two hourly tasks. First, an outside system writes Parquet data to a directory mounted at /mnt/raw_orders/. After this data is written, a Databricks job containing the following code is executed:
Assume that the fields customer_id and order_id serve as a composite key to uniquely identify each order, and that the time field indicates when the record was queued in the source system.
If the upstream system is known to occasionally enqueue duplicate entries for a single order hours apart, which statement is correct?
alexvno
Highly Voted 1 year, 6 months agoIsio05
Highly Voted 1 year agoKadELbied
Most Recent 1 month, 1 week agoa85becd
2 months, 2 weeks agoEZZALDIN
2 months, 3 weeks agoRandomForest
5 months agoarekm
5 months, 2 weeks agobenni_ale
6 months, 4 weeks agocf56faf
7 months, 1 week agoAnanth4Sap
7 months, 3 weeks agom79590530
8 months agoarekm
5 months, 2 weeks agoshaojunni
8 months, 1 week agoquaternion
10 months, 1 week agoQuangTrinh
1 year agoarekm
5 months, 2 weeks ago