問題文
An application must accept a spoken question from the user and reply with spoken audio. Which option correctly identifies the two capabilities required at minimum?
選択肢
- Speech recognition at the input and speech recognition again at the output
- Sentiment analysis to understand the tone, and summarization to shorten the reply
- Speech recognition to turn the question into text, and speech synthesis to turn the reply into audio
- Optical character recognition to read the question, and image generation to draw the reply, because both of these capabilities convert between a signal and text