問題文
A platform team has to serve several models at once: a text classifier saved from one framework, a sentence embedder from another, and a small ranking model. They want one endpoint, one metrics surface, and no separate service per model. Which component addresses that requirement?
選択肢
- An inference server that loads models from several frameworks behind a single interface and exposes their metrics uniformly.
- A tokenizer library, because the tokenizer is the only component that has to be shared across models.
- A notebook environment shared by the team, and a running notebook can hold all the models in memory and answer requests from other services over a local port.
- A vector store, which keeps the model artifacts alongside the document embeddings so that a single lookup can return both the passage and the model that produced it.