問題文
A team wants both citations and the lowest possible latency. Why can these not both be maximized?
選択肢
- Citations increase the max output tokens requirement beyond any model's limit, since every claim in the answer has to carry the passage it came from
- Citations require a provisioned deployment, which is slower because the request has to be routed to the reserved capacity rather than the nearest available instance
- Citations require fine-tuning, which increases per-request latency because the customized weights have to be loaded before the request can be processed
- Producing citations requires a retrieval step whose latency cannot be removed without giving up the citations, so only its speed and count can shrink