問題文
How do you make visual information inside documents reachable by a retrieval-based assistant?
選択肢
- Store the images in blob storage and reference them by URL only, so the assistant can hand the link to the user and let them look at the figure themselves
- Generate descriptions of the figures during ingestion and store them as searchable fields in the index alongside the body text of every one of the documents
- Rely on the model to infer what the figures show from the surrounding text, since a figure in a well-written document is normally explained in the paragraph next to it
- Send the whole document to a multimodal model on every question so the figures are always read from the original file and nothing depends on what was captured at ingestion