問題文
A bank indexes 8,000 policy documents for retrieval. Each document has numbered sections and includes tables of fee schedules. The current pipeline splits every document into fixed 1,000-character chunks. Which two problems does this cause? (Select two.)
選択肢
- Section boundaries are ignored, so a chunk can span the end of one policy rule and the start of an unrelated one, mixing two contexts in one retrieval unit.
- Fixed-length chunking requires re-embedding the entire corpus whenever a single document changes.
- Tables are cut mid-structure, so a retrieved chunk may contain rows without their column headers and the values become uninterpretable.
- Fixed-length chunking always produces more chunks than structure-aware chunking, increasing index size proportionally.
- Fixed-length chunking prevents the use of metadata filters on the index.