問題文
A deployed endpoint answers a single caller quickly but a burst of twenty callers sees long waits and some errors. Which fact about the serving engines explains the waits?
選択肢
- Each engine takes several requests at once and the platform starts a new engine as soon as the queue lengthens.
- Each engine takes one request at a time and the rest of the calls fall back on a cached answer of the caller of the one previous call.
- Each engine takes several requests at once and the platform slows every caller down to keep the memory flat.
- Each engine takes one request at a time and the remaining calls sit in a queue until one of the engines frees up.