問題文
A team must extract organization names and dates from thousands of contracts and store them in a database. They want a library-based pipeline with predictable cost and no generation step. Which choice matches the task?
選択肢
- A quantization toolkit, and reducing the numeric precision of the pipeline is what makes the per-document cost predictable at this scale.
- A natural language processing library that provides tokenization, part-of-speech tagging, and named-entity recognition as pipeline components.
- A guardrails configuration, which validates inputs and outputs but does not extract structured fields from the contracts.
- A vector store, and entity extraction is a retrieval problem and the nearest stored passage for each contract will contain the entities that need extracting.