問題文
A retailer trains a demand model on 900 million rows but serves predictions for at most 300 rows per request from a web application, with a 40 millisecond budget. Which split of training and inference matches these constraints?
選択肢
- Train and serve entirely with SparkML, submitting one Spark job per incoming request, one job for each caller.
- Train with a single-node library on a sample of 2 million rows and serve the same object, because inference latency is the only binding constraint.
- Train with SparkML and serve by writing each request into a Delta table that a streaming job scores every few seconds on a timer.
- Train with SparkML on the cluster, then log a single-node model artifact and serve it from an endpoint, because a 300-row request does not justify starting a distributed job.