問題文
A classifier reaches 0.94 area under the curve but takes 40 milliseconds per request, while the service budget is 15 milliseconds. Which comparison should drive the next decision?
選択肢
- Keep the current model and raise the latency budget, because a model that has already reached this level of quality should not be degraded and the service level agreement is a business decision that can be renegotiated more cheaply than the model can be rebuilt at a smaller size.
- Cache the predictions for the most frequent inputs and accept 40 milliseconds for the rest; the average response time across all requests is what the service budget is really measuring.
- Measure the quality of several smaller configurations against their latency, and pick the point where the quality loss is acceptable within the budget.
- Move inference to a larger device, since latency is determined by device throughput, and a device with more throughput finishes the same request in proportionally less time.